[Skip To Content]

Staff Engineer, Site Reliability Engineering

  • 위치
    • Warren, Michigan
  • 직무 유형 Full time
  • 게시됨
  • Job Requisition JR-202621675

설명

About the role


General Motors is transforming the automotive landscape through its next-generation Software-Defined Vehicle platform. Data is central to that transformation, powering safety, personalization, energy optimization, operational decision-making, and connected customer experiences.

We are seeking a Staff Engineer to help make GM’s data platforms reliable, observable, operable, and scalable. This is a senior technical leadership role for someone who can move comfortably between system-level design, production operations, incident response, automation, and customer partnership.

You will help define and spread the engineering patterns that make services easier to operate. You will work with SRE, data engineering, infrastructure, developer experience, application, and product teams to improve reliability from design through production and continuously improve how the organization operates.


What you’ll do

  • Lead the design and implementation of scalable, fault-tolerant, and observable infrastructure supporting vehicle telemetry, data ingestion, and platform operations.
  • Lead production readiness efforts across multiple teams—engaging directly in code, shaping reliability standards, guiding architectural improvements, and ensuring applications launch with resilient deployments, strong observability, and predictable operations.
  • Design, implement, and improve CI/CD delivery pipelines that make releases repeatable, safe, observable, and fast. Establish appropriate quality gates, artifact promotion, deployment verification, progressive delivery, and rollback practices.
  • Partner across SRE, product, and application teams to design and implement meaningful SLOs, SLIs, observability, monitoring and alerting, runbooks, and operational best practices.
  • Build and improve reusable AI workflows, skills, and evaluations. Apply appropriate validation techniques, including regression testing, structured evaluations, and LLM-as-a-judge approaches where useful.
  • Automate operational work, including incident intake, triage, diagnostics, remediation, evidence collection, service requests, and customer-facing status workflows.
  • Participate in a weekly on-call rotation with 12-hour shifts; the rotation cycles every eight weeks. Lead incident response, communicate clearly under pressure, and coordinate effective mitigation and recovery.
  • Participate in post-incident reviews and drive durable, system-level fixes that prevent recurrence rather than relying on short-term patches or repeated manual workarounds.
  • Partner directly with internal customers to understand their needs, explain technical trade-offs, and improve service outcomes with tact, empathy, and clear communication.
  • Influence technical direction across teams, mentor engineers, and raise engineering standards through design reviews, code reviews, documentation, and hands-on leadership with cross-functional engineering projects.
  • Balance reliability, performance, security, delivery speed, and cost when making technical decisions—especially under pressure.

What you bring

  • 8+ years in SRE, DevOps, or systems engineering, including experience managing or mentoring high-impact teams.
  • Track record of designing and building and maintaining high-scale, cloud-native systems in production (preferably Azure, AWS, or GCP).
  • Hands-on experience architecting observability patterns, including standardized instrumentation, OTEL collector configuration, SLO/SLI definitions, and deploying observability resources like monitors, alerts, and dashboards.
  • Strong understanding of production readiness, service ownership, SLOs, incident management, post-incident learning, and continuous reliability improvement.
  • Experience participating in an on-call rotation and leading technical response to production incidents.
  • Experience designing, operating, and improving CI/CD pipelines. Understanding of GitOps, release strategies, quality gates, deployment verification, progressive delivery, and safe rollback is expected.
  • Strong programming ability in Python, Go, Java, or a comparable language, with disciplined code review, version control, testing, and maintainability practices.
  • Familiarity with AI-assisted software development, LLM application practices, agentic workflows, and evaluation techniques.
  • Ability to work effectively with internal customers, including in difficult or high-pressure situations, with professionalism, tact, and empathy.
  • Excellent ownership attitude and the ability to operate with pace, judgment, and accountability in a high-velocity environment.
  • Ability to influence without relying on formal authority and to increase adoption of shared engineering patterns.
  • Strong written and verbal communication skills for both technical and non-technical audiences.
  • BS / MS / PhD in computer science, engineering, physics, mathematics, or another relevant, technical field

Preferred experience

  • Azure Databricks
  • Azure Event Hubs
  • Azure Kubernetes Service (AKS)
  • Kubernetes configuration management with Helm and Kustomize
  • Infrastructure as code, especially Terraform
  • GitHub Actions, Argo CD, and GitOps-based deployment models
  • Prometheus, Grafana, Datadog, OpenTelemetry, or comparable observability platforms and tools
  • LLM application development and testing with Promptfoo, agentic workflows and reusable AI skills with CoPilot
  • Experience operating large-scale data ingestion, processing, and delivery systems such as Fivetran, Apache Flink, Kafka, and Pulsar
  • Experience with vehicle telemetry, connected-vehicle platforms, or other high-volume event-driven systems

Why join us?

This is an opportunity to shape the reliability foundations of GM’s next generation of connected, software-defined vehicles. You will work on meaningful systems at global scale while helping define how modern SRE teams use automation, observability, operational excellence, and AI to deliver better outcomes for customers.

GM does not provide immigration-related sponsorship for this role. Do not apply for this role if you will need GM immigration sponsorship now or in the future. This includes direct company sponsorship, entry of GM as the immigration employer of record on a government form, and any work authorization requiring a written submission or other immigration support from the company (e.g., H1-B, OPT, STEM OPT, CPT, TN, J-1, etc.)

이 직무는 하이브리드 직무로 분류됩니다. 즉, 선발된 지원자는 특정 근무지로 주 3일 이상(또는 관리자가 지정한 다른 빈도로) 특정 근무지로 출근해야 합니다.

이 직무는 리로케이션 혜택을 받을 수 없습니다. 모든 리로케이션 관련 비용은 최종선정 된 지원자가 부담해야 합니다.

다양성 정보

General Motors는 법적으로 금지된 차별을 배제하는 것은 물론 포용성과 소속감을 진정으로 장려하는 직장이 되기 위해 노력하고 있습니다. 당사는 다양성이 보장되는 환경에서 직원들이 역량을 발휘하고 우리 고객을 위한 더 좋은 제품을 개발할 수 있다고 믿습니다. 따라서 입사에 관심 있는 사람이 있다면 포지션별 주요 업무와 자격을 확인하고 본인이 보유한 기술과 능력에 부합하는 모든 포지션에 적극적으로 지원하기를 장려합니다. 지원자는 채용 과정에서 역할 관련 평가(해당하는 경우) 및/또는 채용 전 스크리닝을 통과해야 합니다.  자세한 정보는 GM 채용 과정 안내를 참고하십시오.

공평한 취업 기회 선언 (미국)

General Motors는 공평한 기회를 제공하는 고용주임을 자부합니다.  자격을 만족하는 지원자는 인종과 피부색, 성별, 성적 지향, 성별 정체성, 국적, 장애, 재향 군인 보호법 적용 여부와 상관없이 채용 후보로서 심사를 받습니다. 

숙소 (미국 및 캐나다)

General Motors는 장애인을 포함한 모든 구직자들에게 취업 기회를 제공합니다. 구직이나 취업 지원에 도움이 되는 합리적인 숙소가 필요한 경우 [email protected]으로 이메일을 보내시거나 800-865-7580으로 전화주십시오. 이메일에, 귀하가 요청하는 특정한 숙소에 대한 설명과 귀하가 지원하는 직무와 채용 요청서 번호를 포함해주세요.