Description
At General Motors, we're turning today's impossible into tomorrow's standard. Our vehicles already move millions of people every day, and we're building the autonomy that will drive them - L2 through L4, on real roads, at real scale. Making self-driving safe at that scale is one of the hardest AI problems there is.
The Data Scaling team owns the data flywheel for AV foundation model pre-training and SFT. We determine what data the AV needs in order to learn driving behaviors at scale, and we define what data quality means across the loop. The team delivers ML models that move the product up the data scaling curves, turning better data composition into measurably better driving behavior. We work with the very large datasets GM already has and we define the next generation of highest-value datasets GM collects. With each major release we aim to 10x the effective data behind our models: more scale, more diversity, and more value extracted from every example.
Why Join Us?
- Train on driving data almost nobody else has - real-world miles from GM's fleets, plus synthetic sim data - scaling into billions of examples. Then decide which ones are worth it: ten thousand near-identical highway miles teach the model less than one unprotected left turn in the rain. Mixture design, curation, mining, and evaluation are how you find out which is which, working alongside other MLEs and research scientists.
- Work on questions with no textbook answers yet. Scaling laws for language are well mapped by now; for embodied driving data - heavy-tailed, safety-constrained, closed-loop - they aren't. You'd be helping write them, and we support publishing what you find.
- See your results in the world rather than on a leaderboard. The models this team ships change how the vehicle behaves on real roads, and that behavior comes back as the evidence for your next iteration.
As a Senior AI/ML Engineer in the Embodied AI Data Foundations organization, you will be an individual contributor developing data-centric AI solutions that directly improve autonomous driving performance. You will design and run the data curation and model training recipes that produce models capable of safe, reliable behavior across diverse real-world scenarios, drawing on both real and synthetic data.
What You'll Do
- Design and run experiments that connect data composition to model behavior: dataset mixtures, sampling strategies, curricula, and scaling-law studies that tell us where to invest next.
- Apply methods such as self-supervised pre-training, imitation learning, reinforcement learning, and foundation-model fine-tuning to driving behavior, trajectory generation, and perception tasks.
- Develop data curation and mining methods - auto-labeling, deduplication, difficulty and uncertainty estimation, long-tail and out-of-distribution scenario discovery - to raise the value of every training example.
- Define offline metrics and evaluations that actually predict on-road behavior, and use them to make model and data decisions from evidence rather than intuition.
- Trace model failures back to their root cause in the data, then close the loop by specifying the data needed to fix them.
- Train models at scale across large multi-GPU/multi-node datasets, partnering with platform teams on the pipelines and tooling this requires.
- Collaborate with cross-functional teams to bring models into onboard driving systems, and document learnings and best practices along the way.
- Follow relevant literature and bring promising advances into our recipes and evaluations.
Your Skills & Abilities
Required:
- Master's or PhD in Computer Science, Robotics, Machine Learning.
- Strong ML fundamentals: you can design a clean experiment, pick the right baseline, read an ablation, and tell signal from noise.
- Proficiency in Python and PyTorch, with experience training models on large datasets.
- Hands-on experience with data-centric ML: curation, sampling, labeling, or evaluation of large training sets.
- Working knowledge of large-scale foundation models and how they are pre-trained, fine-tuned, and aligned.
- Solid data analysis skills (NumPy, Pandas; SQL or Spark for large datasets).
- Demonstrated ability to deliver applied ML results under real-world constraints and timelines.
- Clear communication: you can explain a result and its limits to both engineers and non-experts.
Preferred:
- PhD, publications, or open-source contributions in representation learning, multimodal or vision-language models, generative models, RL, or data-centric ML.
- Experience with robotics, autonomous driving, or other embodied AI systems.
- Experience with synthetic and simulation data, including sim-to-real transfer.
- Familiarity with production ML deployment workflows.
Remote/Hybrid: This role is categorized as fully remote or hybrid.
Compensation: The compensation information is a good faith estimate only. It is based on what a successful applicant might be paid in accordance with applicable state laws. The compensation may not be representative for positions located outside of the California Bay Area.
- The salary range for this role is $159,300.00 to $230,700.00. The actual base salary a successful candidate will be offered within this range will vary based on factors relevant to the position.
- Bonus Potential: An incentive pay program offers payouts based on company performance, job level, and individual performance.
Benefits:
- GM offers a variety of health and wellbeing benefit programs. Benefit options include medical, dental, vision, Health Savings Account, Flexible Spending Accounts, retirement savings plan, sickness and accident benefits, life insurance, paid vacation & holidays, tuition assistance programs, employee assistance program, GM vehicle discounts and more.
This job may be eligible for relocation benefits.
#LI-CX1
#GM-AV-1
Renseignements sur la diversité
General Motors est résolue à être un lieu de travail qui est non seulement exempt de discrimination illégale, mais aussi un endroit qui favorise véritablement l'inclusion et l'appartenance. Nous sommes convaincus que la diversité de la main-d'œuvre permet de créer un environnement dans lequel nos employés peuvent s'épanouir et développer de meilleurs produits pour nos clients. Nous encourageons les candidats intéressés à consulter les principales responsabilités et compétences requises pour chaque rôle et à postuler à tout poste qui leur correspond. Dans le cadre du processus de recrutement, les candidats peuvent devoir, le cas échéant, réussir une évaluation liée au poste ou une présélection d'emploi avant d'être embauchés. Pour en savoir plus, consultez notre processus de recrutement.
Déclaration concernant l'égalité d'accès à l'emploi (É.-U.)
General Motors est fière d'être un employeur souscrivant au principe de l'égalité d'accès à l'emploi. Tous les candidats qualifiés seront pris en compte, sans égard à la race, à la couleur, à la religion, au sexe, à l'orientation sexuelle, à l'identité de genre, à l'origine ethnique, aux situations de handicap ou au statut protégé d'ancien combattant.
Aménagements (É.-U. et Canada)
General Motors offre des occasions à tous les chercheurs d'emploi, y compris les personnes handicapées. Si vous avez besoin d'un accommodement raisonnable pour vous aider dans votre recherche d'emploi ou la soumission de votre candidature, envoyez-nous un courriel à l'adresse [email protected] ou appelez-nous au 800 865-7580. Veuillez inclure dans votre courriel une description spécifique du type d'accommodement demandé, ainsi que le titre d'emploi et le numéro de demande du poste auquel vous postulez.
