[Skip To Content]

Staff Engineer, Site Reliability Engineering

  • [Location]
    • Markham, Ontario
  • 职位类型 Full time
  • 发表
  • Job Requisition JR-202618328

描述

Vacancy Status:

Yes - This posting is for an existing vacancy within the organization and is open to new applications. (Backfill)

AI Disclosure:

As part of the application process, Artificial Intelligence will be used in the hiring process for this role

Work Arrangement: Hybrid: This role is categorized as hybrid. This means the successful candidate is expected to report to Markham office three times per week, at minimum.

About the role  

General Motors is transforming the automotive landscape through its next-generation Software-Defined Vehicle platform. Data is central to that transformation, powering safety, personalization, energy optimization, operational decision-making, and connected customer experiences. 

We are seeking a Staff Engineer to help make GM’s data platforms reliable, observable, operable, and scalable. This is a senior technical leadership role for someone who can move comfortably between system-level design, production operations, incident response, automation, and customer partnership. 

You will help define and spread the engineering patterns that make services easier to operate. You will work with SRE, data engineering, infrastructure, developer experience, application, and product teams to improve reliability from design through production and continuously improve how the organization operates. 

What you’ll do

  • Lead the design and implementation of scalable, fault-tolerant, and observable infrastructure supporting vehicle telemetry, data ingestion, and platform operations. 

  • Lead production readiness efforts across multiple teams—engaging directly in code, shaping reliability standards, guiding architectural improvements, and ensuring applications launch with resilient deployments, strong observability, and predictable operations. 

  • Design, implement, and improve CI/CD delivery pipelines that make releases repeatable, safe, observable, and fast. Establish appropriate quality gates, artifact promotion, deployment verification, progressive delivery, and rollback practices. 

  • Partner across SRE, product, and application teams to design and implement meaningful SLOs, SLIs, observability, monitoring and alerting, runbooks, and operational best practices. 

  • Build and improve reusable AI workflows, skills, and evaluations. Apply appropriate validation techniques, including regression testing, structured evaluations, and LLM-as-a-judge approaches where useful. 

  • Automate operational work, including incident intake, triage, diagnostics, remediation, evidence collection, service requests, and customer-facing status workflows. 

  • Participate in a weekly on-call rotation with 12-hour shifts; the rotation cycles every eight weeks. Lead incident response, communicate clearly under pressure, and coordinate effective mitigation and recovery. 

  • Participate in post-incident reviews and drive durable, system-level fixes that prevent recurrence rather than relying on short-term patches or repeated manual workarounds. 

  • Partner directly with internal customers to understand their needs, explain technical trade-offs, and improve service outcomes with tact, empathy, and clear communication. 

  • Influence technical direction across teams, mentor engineers, and raise engineering standards through design reviews, code reviews, documentation, and hands-on leadership with cross-functional engineering projects. 

  • Balance reliability, performance, security, delivery speed, and cost when making technical decisions—especially under pressure. 
     

What you bring  

  • 8+ years in SRE, DevOps, or systems engineering, including experience managing or mentoring high-impact teams. 

  • Track record of designing and building and maintaining high-scale, cloud-native systems in production (preferably Azure, AWS, or GCP). 

  • Hands-on experience architecting observability patterns, including standardized instrumentation, OTEL collector configuration, SLO/SLI definitions, and deploying observability resources like monitors, alerts, and dashboards. 

  • Strong understanding of production readiness, service ownership, SLOs, incident management, post-incident learning, and continuous reliability improvement. 

  • Experience participating in an on-call rotation and leading technical response to production incidents. 

  • Experience designing, operating, and improving CI/CD pipelines. Understanding of GitOps, release strategies, quality gates, deployment verification, progressive delivery, and safe rollback is expected. 

  • Strong programming ability in Python, Go, Java, or a comparable language, with disciplined code review, version control, testing, and maintainability practices. 

  • Familiarity with AI-assisted software development, LLM application practices, agentic workflows, and evaluation techniques. 

  • Ability to work effectively with internal customers, including in difficult or high-pressure situations, with professionalism, tact, and empathy. 

  • Excellent ownership attitude and the ability to operate with pace, judgment, and accountability in a high-velocity environment. 

  • Ability to influence without relying on formal authority and to increase adoption of shared engineering patterns. 

  • Strong written and verbal communication skills for both technical and non-technical audiences. 

  • BS / MS / PhD in computer science, engineering, physics, mathematics, or another relevant, technical field 

Preferred experience  

  • Azure Databricks 

  • Azure Event Hubs 

  • Azure Kubernetes Service (AKS) 

  • Kubernetes configuration management with Helm and Kustomize 

  • Infrastructure as code, especially Terraform 

  • GitHub Actions, Argo CD, and GitOps-based deployment models 

  • Prometheus, Grafana, Datadog, OpenTelemetry, or comparable observability platforms and tools 

  • LLM application development and testing with Promptfoo, agentic workflows and reusable AI skills with CoPilot 

  • Experience operating large-scale data ingestion, processing, and delivery systems such as Fivetran, Apache Flink, Kafka, and Pulsar 

  • Experience with vehicle telemetry, connected-vehicle platforms, or other high-volume event-driven systems 

Why Join Us?

This is more than an engineering role — it’s an opportunity to shape the future of mobility . At GM, you’ll join a team committed to cutting-edge technology, sustainability , and inclusive innovation . With meaningful projects, a collaborative culture, and a global mission, your impact will be tangible and far-reaching.

Compensation: 

The salary range for this role is $147,000 to $196,600. The actual base salary a successful candidate will be offered within this range will vary based on factors relevant to the position.

GM DOES NOT PROVIDE IMMIGRATION-RELATED SPONSORSHIP FOR THIS ROLE. DO NOT APPLY FOR THIS ROLE IF YOU WILL NEED GM IMMIGRATION SPONSORSHIP NOW OR IN THE FUTURE.

Benefits Overview

The goal of the General Motors of Canada total rewards program is to support the health and well-being of you and your family. Our comprehensive compensation plan currently includes the following benefits, in addition to many others:

  • Paid time off including vacation days, holidays, and supplemental benefits for pregnancy, parental and adoption leave;
  • Healthcare, dental, and vision benefits;
  • Life insurance plans to cover you and your family;
  • Company and matching contributions to a Defined Contribution Pension plan to help you save for retirement;
  • GM Vehicle Purchase Plan for you and your family.

Renseignements sur la diversité

General Motors est résolue à être un lieu de travail qui est non seulement exempt de discrimination illégale, mais aussi un endroit qui favorise véritablement l'inclusion et l'appartenance. Nous sommes convaincus que la diversité de la main-d'œuvre permet de créer un environnement dans lequel nos employés peuvent s'épanouir et développer de meilleurs produits pour nos clients. Nous encourageons les candidats intéressés à consulter les principales responsabilités et compétences requises pour chaque rôle et à postuler à tout poste qui leur correspond. Dans le cadre du processus de recrutement, les candidats peuvent devoir, le cas échéant, réussir une évaluation liée au poste ou une présélection d'emploi avant d'être embauchés.  Pour en savoir plus, consultez notre processus de recrutement.

Déclaration concernant l'égalité d'accès à l'emploi (É.-U.)

General Motors est fière d'être un employeur souscrivant au principe de l'égalité d'accès à l'emploi.  Tous les candidats qualifiés seront pris en compte, sans égard à la race, à la couleur, à la religion, au sexe, à l'orientation sexuelle, à l'identité de genre, à l'origine ethnique, aux situations de handicap ou au statut protégé d'ancien combattant. 

Aménagements (É.-U. et Canada)

General Motors offre des occasions à tous les chercheurs d'emploi, y compris les personnes handicapées. Si vous avez besoin d'un accommodement raisonnable pour vous aider dans votre recherche d'emploi ou la soumission de votre candidature, envoyez-nous un courriel à l'adresse [email protected] ou appelez-nous au 800 865-7580. Veuillez inclure dans votre courriel une description spécifique du type d'accommodement demandé, ainsi que le titre d'emploi et le numéro de demande du poste auquel vous postulez.