AIML - Machine Learning Research Lead, RL Agents, MLR

Apple Inc.

Cupertino, Northern (CA, KY)

Hybrid

USD 216,000 - 394,000

Full time

5 days ago
Be an early applicant
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Benefits offered by this job

Apple stock programs
Medical coverage
Retirement benefits
Tuition reimbursement

Job summary

Apple Inc. in Cupertino, CA is seeking a hands-on research lead to drive reinforcement learning and post-training work for agentic AI, and to manage a small team of senior researchers in RL, agentic tool-calling, synthetic data generation, and multimodal action models.

You will help set direction for infrastructure, training, runtime and evaluation for interactive agents and on-device/hybrid deployment. The role combines research leadership with hands-on coding and experimentation, publishing in

Qualifications

  • PhD in machine learning or a related field, or equivalent research experience.
  • 7-10+ years of research experience beyond PhD in industry or as an academic research lead.
  • Strong track record in RL and/or post-training of large models, demonstrated through publications, open-source contributions, or shipped systems.
  • Leadership experience: setting and defending a research direction over multiple years, and directing others' work — through direct reports, PhD students, postdocs, or sustained project teams.
  • Experience owning ML infrastructure, frameworks and codebases, including open-source research frameworks or environment suites others build on.

Responsibilities

  • Lead research on RL and post-training for agentic capabilities: reward, preference optimization, and verifier design, training recipes, and evaluation for tool calling, coding, and multi-step interactive tasks.
  • Build and own synthetic data and task-generation pipelines — generating diverse, verifiable tasks and environments, along with the interactive environments and benchmarks that go with them, and the curricula that turn them into capable agents.
  • Drive codebases and infrastructure for the core RL research effort and help engage partner teams to use and co-develop the framework.
  • Manage and mentor a small team (3–4) of senior researchers and research engineers with distinct specialties, shaping a shared research direction while protecting room for bottom-up, idea-driven work.
  • Stay hands-on: run experiments, write code, and contribute directly to the team's most important technical problems.
  • Connect post-training research to efficiency and deployment: what works under on-device and hybrid compute constraints, and how method design interacts with hardware.
  • Collaborate across the organization on adjacent directions, including methods for environment and agent co-optimization, self-improvement, world models used as planners or policies, and personalized long-context agents.
  • Publish in top venues and engage with the broader research community.

Skills

Reinforcement Learning
Post-training methods
Leadership
ML infrastructure

Education

PhD in ML or related field

Tools

TensorFlow
PyTorch
JAX
OpenAI Gym

Job description

AIML - Machine Learning Research Lead, RL Agents, MLR

Cupertino, California, United States Machine Learning and AI

We are looking for a hands‑on research lead to drive our work on reinforcement learning and post‑training for agentic AI, and to manage a small team of senior researchers working on related problems in RL, agentic tool‑calling, synthetic environment generation, model scaling, and multimodal action models. You will help set direction for how we develop infrastructure, training, runtime and evaluation procedures for interactive agents — tool calling, coding, computer use, and long‑horizon tasks. This role sits inside a research organization pursuing first‑principles approaches to core AI problems: generative foundation models across modalities (text, images, graphs, scientific and engineering data), vision‑language modeling and implicit world modeling, self‑supervised learning, and search and evolutionary methods for optimizing both agents and the environments they learn in. A distinctive part of our agenda is designing methods that fit Apple’s deployment reality — on‑device and hybrid (device plus private cloud) execution, co‑designed with current and future hardware — and that take advantage of what this ecosystem uniquely enables, such as deeply personalized, long‑context agentic experiences. We aim for both field‑changing research and direct impact on Apple products and internal engineering processes. MLR is a research group first. Management here is about spreading the load of running a team, not stepping away from the work — everyone, including leads, stays hands‑on. We support continued engagement with the academic community: publishing, conference service, student collaboration, and internships.

Description
  • Lead research on RL and post‑training for agentic capabilities: reward, preference optimization, and verifier design, training recipes, and evaluation for tool calling, coding, and multi‑step interactive tasks.
  • Build and own synthetic data and task‑generation pipelines — generating diverse, verifiable tasks and environments, along with the interactive environments and benchmarks that go with them, and the curricula that turn them into capable agents.
  • Drive codebases and infrastructure for the core RL research effort and help engage partner teams to use and co‑develop the framework.
  • Manage and mentor a small team (3–4) of senior researchers and research engineers with distinct specialties, shaping a shared research direction while protecting room for bottom‑up, idea‑driven work.
  • Stay hands‑on: run experiments, write code, and contribute directly to the team's most important technical problems.
  • Connect post‑training research to efficiency and deployment: what works under on‑device and hybrid compute constraints, and how method design interacts with hardware.
  • Collaborate across the organization on adjacent directions, including methods for environment and agent co‑optimization, self‑improvement, world models used as planners or policies, and personalized long‑context agents.
  • Publish in top venues and engage with the broader research community.
Minimum Qualifications
  • PhD in machine learning or a related field, or equivalent research experience
  • 7-10+ years of research experience beyond PhD in industry or as an academic research lead
  • Strong track record in RL and/or post‑training of large models, demonstrated through publications, open‑source contributions, or shipped systems
  • Leadership experience: setting and defending a research direction over multiple years, and directing others' work — through direct reports, PhD students, postdocs, or sustained project teams. Formal management experience is welcome but not required
  • Experience owning ML infrastructure, frameworks and codebases, including open‑source research frameworks or environment suites others build on
Preferred Qualifications
  • Experience taking research from idea to product or production impact
  • Familiarity with efficiency‑aware modeling: small models, mixture‑of‑experts, quantization, distillation, inference‑cost constraints, or hardware‑aware method design
  • Interest or background in open‑endedness, evolutionary computation, curriculum or environment design, multi‑agent systems, or self‑improving systems
  • Principled or theoretical grounding in RL — representation, exploration, or optimization views of policy learning — alongside strong empirical work
  • Breadth across core machine learning — generative models, self‑supervised learning, pre‑training — and perspective on the field's longer arcs, not only its most recent methods
  • Experience growing other researchers, and managing researchers and engineers with heterogeneous specialties and synthesizing their work toward a common goal
  • Experience owning a large RL or post‑training codebase

At Apple, base pay is one part of our total compensation package and is determined within a range. This provides the opportunity to progress as you grow and develop within a role. The base pay range for this role is between $216,200 and $394,000, and your base pay will depend on your skills, qualifications, experience, and location. Apple employees also have the opportunity to become an Apple shareholder through participation in Apple’s discretionary employee stock programs. Apple employees are eligible for discretionary restricted stock unit awards, and can purchase Apple stock at a discount if voluntarily participating in Apple’s Employee Stock Purchase Plan. You’ll also receive benefits including: Comprehensive medical and dental coverage, retirement benefits, a range of discounted products and free services, and for formal education related to advancing your career at Apple, reimbursement for certain educational expenses — including tuition. Additionally, this role might be eligible for discretionary bonuses or commission payments as well as relocation. Learn more about Apple Benefits. Note: Apple benefit, compensation and employee stock programs are subject to eligibility requirements and other terms of the applicable plan or program.

Apple is an equal opportunity employer that is committed to inclusion and diversity. We seek to promote equal opportunity for all applicants without regard to race, color, religion, sex, sexual orientation, gender identity, national origin, disability, Veteran status, or other legally protected characteristics. Learn more about your EEO rights as an applicant

At Apple, we believe accessibility is a fundamental human right. You’ll find that idea reflected in everything here — in our culture, our benefits and our digital tools. By welcoming as many perspectives as possible, we help you build a career where you feel like you belong. Learn about accessibility in Apple’s workplace. Learn about reasonable accommodations for job applicants

Apple accepts applications to this posting on an ongoing basis.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

AIML - Machine Learning Research Lead, RL Agents, MLR
AIML - Machine Learning Research Lead, RL Agents, MLR

Socket.dev • Cupertino (CA)

On-site
USD 220,000 - 320,000
ML Agent Engineer
ML Agent Engineer

Apple Inc. • Seattle (WA)

On-site
USD 150,000 - 278,000
Employee stock options
Relocation assistance
Education reimbursement
Research Scientist/ Engineer - Agentic Workflows
Research Scientist/ Engineer - Agentic Workflows

Apple Inc. • Santa Clara (CA), Northern (KY)

Hybrid
USD 185,000 - 325,000
AIML - Staff Machine Learning Engineer
AIML - Staff Machine Learning Engineer

Apple Inc. • Seattle (WA)

On-site
USD 139,000 - 259,000
Comprehensive medical and dental coverage
Retirement benefits
Discounted products and services
+2
AIML – Machine Learning Researcher, MLR
AIML – Machine Learning Researcher, MLR

NLP PEOPLE • Cambridge (MA)

On-site
USD 166,000 - 250,000
Medical and dental coverage
Retirement benefits
Employee stock programs
+2
AIML - Machine Learning Researcher, Data and ML Innovation
AIML - Machine Learning Researcher, Data and ML Innovation

Apple Inc. • Santa Clara (CA)

On-site
USD 181,000 - 319,000
Comprehensive medical and dental coverage
Retirement benefits
Discounted products and free services
+1
Staff Research Scientist, Siri Innovation Studio
Staff Research Scientist, Siri Innovation Studio

Apple Inc. • Cupertino (CA), Northern (KY)

Hybrid
USD 216,000 - 325,000
Medical and dental coverage
Stock programs
Tuition reimbursement
Research Scientist/ Engineer - Agentic Workflows
Research Scientist/ Engineer - Agentic Workflows

Apple Inc. • Seattle (WA)

On-site
USD 175,000 - 309,000
Machine Learning Researcher Multi-Modal Reasoning
Machine Learning Researcher Multi-Modal Reasoning

Apple Inc. • Cupertino (CA)

On-site
USD 185,000 - 325,000
Comprehensive medical and dental
Retirement benefits
Discounted Apple products
+3
AI/ML Engineer: System RF Data Ecosystem
AI/ML Engineer: System RF Data Ecosystem

Apple Inc. • Cupertino (CA), Northern (KY)

Hybrid
USD 185,000 - 325,000
Comprehensive medical and dental cover
Employee Stock Purchase Plan
Discretionary bonuses
+2