Descrição
Work Arrangement:
Hybrid: This internship is categorized as on-site. The selected intern is expected to report to the office 3 days a week.
Location:
Sunnyvale, CA
About the Team
The Data Scaling team builds the data, machine learning, and infrastructure foundations that enable Embodied AI models to improve with scale. We work across data collection, curation, mining, labeling, dataset quality, model inputs, distributed training, experiment workflows, and the systems that help researchers and engineers develop, evaluate, and deploy models more efficiently.
Our work supports large-scale autonomous driving models and the broader Embodied AI flywheel. We combine machine learning, data engineering, distributed systems, and software engineering to make high-quality, product-aligned driving data available for perception, planning, trajectory generation, and other autonomy capabilities.
About the Role
As an Embodied AI Data Scaling Intern, you will work on a well-scoped project that improves the scale, quality, efficiency, or reliability of data and model development for autonomous driving. You will collaborate with researchers and engineers to build data pipelines, analyze large datasets, improve training workflows, or develop infrastructure that increases the number and quality of experiments the team can run.
Potential focus areas include:
- Data Curation and Quality: Build methods to select, filter, balance, and validate large driving datasets aligned with product and modeling needs.
- Data Mining and Scenario Discovery: Identify valuable, rare, or challenging driving situations and develop tools to improve coverage of long-tail scenarios.
- Dataset and Labeling Pipelines: Improve automated labeling, data transformation, feature generation, and dataset release workflows.
- Distributed Training and Model Scaling: Optimize the systems, input pipelines, and compute workflows used to train large models on large datasets.
- ML Experimentation and Flywheel Infrastructure: Build tools that help teams develop, train, evaluate, debug, and deploy models more quickly and reliably.
What You’ll Do
- Develop data pipelines and tooling for large-scale, multimodal autonomous driving datasets.
- Analyze data quality, coverage, distribution, and performance impact using quantitative methods.
- Prototype and evaluate approaches for data mining, curation, labeling, sampling, or scenario discovery.
- Improve training throughput, data loading, experiment reproducibility, or resource utilization in distributed computing environments.
- Collaborate with machine learning researchers, data engineers, infrastructure engineers, and autonomy teams.
- Build visualizations, metrics, dashboards, and evaluation workflows to communicate data and model behavior.
- Contribute to production-quality software through design reviews, code reviews, automated testing, continuous integration, and documentation.
- Present technical findings and document experiments, results, and recommendations.
Required Qualifications
- Currently enrolled in or pursuing a Master’s or Ph.D. degree in Computer Science, Machine Learning, Data Science, Electrical Engineering, or a related technical field.
- Demonstrated experience through coursework, research, academic projects, or professional work in machine learning, data engineering, distributed systems, or a related area.
- Strong programming skills in Python.
- Experience working with data processing, databases, machine learning pipelines, or large-scale datasets.
- Strong analytical and problem-solving skills, with the ability to use quantitative analysis to guide decisions.
- Ability to work collaboratively in a cross-functional, team-oriented environment.
- Strong written, verbal, and presentation skills.
- Availability to work full-time, 40 hours per week, during the internship period.
Preferred Qualifications
- Experience with PyTorch, TensorFlow, JAX, or another machine learning framework.
- Experience with distributed training, parallel computing, high-performance computing, cloud infrastructure, or GPU-based workflows.
- Familiarity with multimodal sensor data, autonomous vehicles, robotics, computer vision, or large-scale time-series data.
- Experience with SQL, data warehouses, workflow orchestration, data quality systems, or distributed data-processing frameworks.
- Familiarity with foundation models, self-supervised learning, imitation learning, reinforcement learning, or deep learning.
- Experience with C++ or another systems programming language.
- Must be graduating between December 2027 and June 2028
- Intent to return to a degree program after completion of the internship.
Compensation:
- The monthly salary range for this role is $11,800 – $14,600 per month
- GM will provide a one-time lump sum taxable stipend payment to eligible students selected for the 2027 Student Program.
What You’ll Get from Us
- Paid U.S. GM holidays.
- GM Family First Vehicle Discount Program.
- Potential for growth within GM based on performance and business needs.
- Intern events and opportunities to network with company leaders and peers.
- Mentorship and hands-on experience with the data and ML foundations behind autonomy
Informações sobre diversidade
A General Motors está comprometida em ser um local de trabalho que não só é livre de discriminação ilegal, como estimula verdadeiramente a inclusão e integração. Acreditamos enfaticamente que a diversidade na força de trabalho cria um ambiente no qual nossos colaboradores podem crescer e desenvolver melhores produtos para nossos clientes. Incentivamos os candidatos interessados a analisar as principais responsabilidades e qualificações de cada função e a se candidatar a qualquer cargo que corresponda a suas habilidades e capacidades. Os candidatos no processo de recrutamento podem, quando aplicável, ser solicitados a concluir com sucesso uma ou mais avaliações relacionadas à função e/ou uma seleção pré-emprego antes de iniciar o emprego. Para saber mais, acesse Como contratamos.
Declaração de Igualdade de Oportunidades de Emprego (EUA)
A General Motors tem orgulho de ser um empregador que oferece oportunidades iguais. Todos os candidatos qualificados serão considerados para o emprego, independentemente de raça, cor, religião, sexo, orientação sexual, identidade de gênero, origem nacional, deficiência ou status como veterano protegido.
Adaptações (EUA e Canadá)
A General Motors oferece oportunidades a todos os candidatos a emprego, incluindo pessoas com deficiências. Se você precisa de uma adaptação razoável para ajudá-lo na sua pesquisa de cargos ou solicitação de emprego, fale conosco pelo e-mail [email protected] ou pelo telefone 800-865-7580. No seu e-mail, inclua uma descrição da adaptação específica que você está solicitando assim como o nome do cargo e o número de requisição do cargo ao qual está se candidatando.
