Descripción
At General Motors, we're turning today's impossible into tomorrow's standard. Our vehicles already move millions of people every day, and we're building the autonomy that will drive them - L2 through L4, on real roads, at real scale. Making self-driving safe at that scale is one of the hardest AI problems there is.
The Data Scaling team owns the data flywheel for AV foundation model pre-training and SFT. We determine what data the AV needs in order to learn driving behaviors at scale, and we define what data quality means across the loop. The team delivers ML models that move the product up the data scaling curves, turning better data composition into measurably better driving behavior. We work with the very large datasets GM already has and we define the next generation of highest-value datasets GM collects. With each major release we aim to 10x the effective data behind our models: more scale, more diversity, and more value extracted from every example.
Why Join Us?
- Train on driving data almost nobody else has - real-world miles from GM's fleets, plus synthetic sim data - scaling into billions of examples. Then decide which ones are worth it: ten thousand near-identical highway miles teach the model less than one unprotected left turn in the rain. Mixture design, curation, mining, and evaluation are how you find out which is which, working alongside other MLEs and research scientists.
- Work on questions with no textbook answers yet. Scaling laws for language are well mapped by now; for embodied driving data - heavy-tailed, safety-constrained, closed-loop - they aren't. You'd be helping write them, and we support publishing what you find.
- See your results in the world rather than on a leaderboard. The models this team ships change how the vehicle behaves on real roads, and that behavior comes back as the evidence for your next iteration.
As a Senior AI/ML Engineer in the Embodied AI Data Foundations organization, you will be an individual contributor developing data-centric AI solutions that directly improve autonomous driving performance. You will design and run the data curation and model training recipes that produce models capable of safe, reliable behavior across diverse real-world scenarios, drawing on both real and synthetic data.
What You'll Do
- Design and run experiments that connect data composition to model behavior: dataset mixtures, sampling strategies, curricula, and scaling-law studies that tell us where to invest next.
- Apply methods such as self-supervised pre-training, imitation learning, reinforcement learning, and foundation-model fine-tuning to driving behavior, trajectory generation, and perception tasks.
- Develop data curation and mining methods - auto-labeling, deduplication, difficulty and uncertainty estimation, long-tail and out-of-distribution scenario discovery - to raise the value of every training example.
- Define offline metrics and evaluations that actually predict on-road behavior, and use them to make model and data decisions from evidence rather than intuition.
- Trace model failures back to their root cause in the data, then close the loop by specifying the data needed to fix them.
- Train models at scale across large multi-GPU/multi-node datasets, partnering with platform teams on the pipelines and tooling this requires.
- Collaborate with cross-functional teams to bring models into onboard driving systems, and document learnings and best practices along the way.
- Follow relevant literature and bring promising advances into our recipes and evaluations.
Your Skills & Abilities
Required:
- Master's or PhD in Computer Science, Robotics, Machine Learning.
- Strong ML fundamentals: you can design a clean experiment, pick the right baseline, read an ablation, and tell signal from noise.
- Proficiency in Python and PyTorch, with experience training models on large datasets.
- Hands-on experience with data-centric ML: curation, sampling, labeling, or evaluation of large training sets.
- Working knowledge of large-scale foundation models and how they are pre-trained, fine-tuned, and aligned.
- Solid data analysis skills (NumPy, Pandas; SQL or Spark for large datasets).
- Demonstrated ability to deliver applied ML results under real-world constraints and timelines.
- Clear communication: you can explain a result and its limits to both engineers and non-experts.
Preferred:
- PhD, publications, or open-source contributions in representation learning, multimodal or vision-language models, generative models, RL, or data-centric ML.
- Experience with robotics, autonomous driving, or other embodied AI systems.
- Experience with synthetic and simulation data, including sim-to-real transfer.
- Familiarity with production ML deployment workflows.
Remote/Hybrid: This role is categorized as fully remote or hybrid.
Compensation: The compensation information is a good faith estimate only. It is based on what a successful applicant might be paid in accordance with applicable state laws. The compensation may not be representative for positions located outside of the California Bay Area.
- The salary range for this role is $159,300.00 to $230,700.00. The actual base salary a successful candidate will be offered within this range will vary based on factors relevant to the position.
- Bonus Potential: An incentive pay program offers payouts based on company performance, job level, and individual performance.
Benefits:
- GM offers a variety of health and wellbeing benefit programs. Benefit options include medical, dental, vision, Health Savings Account, Flexible Spending Accounts, retirement savings plan, sickness and accident benefits, life insurance, paid vacation & holidays, tuition assistance programs, employee assistance program, GM vehicle discounts and more.
This job may be eligible for relocation benefits.
#LI-CX1
#GM-AV-1
Información sobre diversidad
General Motors se compromete a ser un lugar de trabajo en el cual no solo no haya discriminación indebida, sino que fomente con sinceridad la inclusión y el sentido de pertenencia. Creemos firmemente que la diversidad del personal crea un entorno en el cual nuestros empleados pueden prosperar y desarrollar mejores productos para nuestros clientes. Instamos a los candidatos interesados a que revisen las responsabilidades y aptitudes clave para cada puesto y se postulen para los puestos que coincidan con sus habilidades y capacidades. Es posible que, cuando corresponda, se les pida a los solicitantes que están en el proceso de contratación que completen satisfactoriamente una o más evaluaciones relacionadas con su función y/o una evaluación previa al empleo antes de comenzar a trabajar. Para obtener más información, visite Cómo contratamos.
Declaración de igualdad de oportunidades en el empleo (EE.UU.)
General Motors se enorgullece de ser un empleador que ofrece igualdad de oportunidades. Todos los solicitantes calificados serán tenidos en cuenta para el empleo sin distinción de raza, color, religión, sexo, orientación sexual, identidad de género, nacionalidad, discapacidad o condición de veterano protegido.
Adecuaciones (EE.UU. y Canadá)
General Motors ofrece oportunidades a todos los solicitantes de empleo, incluyendo las personas con discapacidades. Si necesita una adecuación razonable para ayudarle con su búsqueda o solicitud de empleo, envíenos un correo electrónico a [email protected] o llámenos al 800-865-7580. En su correo electrónico, incluya una descripción del puesto específico que está solicitando, así como el título del empleo y el número de solicitud del puesto que está solicitando.




