AIML - Machine Learning Research Lead, RL Agents, MLR

Socket.dev

Cupertino (CA)

On-site

USD 220,000 - 320,000

Full time

6 days ago
Be an early applicant
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Job summary

Socket.dev is seeking a hands-on research lead to guide reinforcement learning and post-training initiatives for agentic AI, while managing a small team of senior researchers across RL, tool-calling, synthetic environments, and multimodal action models.

You will shape the infrastructure, training, runtime, and evaluation strategies for interactive agents, and remain actively involved in experiments, coding, and publishing.

Qualifications

  • PhD in machine learning or related field, or equivalent research experience.
  • 7-10+ years of research experience beyond PhD in industry or academia.
  • Strong track record in RL and/or post-training of large models, evidenced by publications, open-source contributions, or shipped systems.
  • Leadership experience coordinating researchers and engineers over multiple years.

Responsibilities

  • Lead RL and post-training research for agentic capabilities, including reward, preference optimization, and verifier design.
  • Build and own synthetic data and task-generation pipelines and the environments and benchmarks that accompany them.
  • Drive codebases and infrastructure for the core RL research effort and coordinate with partner teams.
  • Manage and mentor a small team (3–4) of senior researchers and research engineers, shaping a shared direction while preserving bottom-up work.
  • Stay hands-on: run experiments, write code, and contribute directly to the team's core technical problems.
  • Bridge post-training research with deployment considerations for on-device and hybrid compute environments.
  • Collaborate across the organization on environment-agent co-optimization, self-improvement, and long-context agents.
  • Publish in top venues and engage with the broader research community.

Skills

RL research
Leadership
ML infrastructure
Hands-on engineering

Education

PhD or equivalent

Tools

PyTorch
RL frameworks

Job description

We are looking for a hands-on research lead to drive our work on reinforcement learning and post-training for agentic AI, and to manage a small team of senior researchers working on related problems in RL, agentic tool-calling, synthetic environment generation, model scaling, and multimodal action models. You will help set direction for how we develop infrastructure, training, runtime and evaluation procedures for interactive agents — tool calling, coding, computer use, and long-horizon tasks. This role sits inside a research organization pursuing first-principles approaches to core AI problems: generative foundation models across modalities (text, images, graphs, scientific and engineering data), vision-language modeling and implicit world modeling, self-supervised learning, and search and evolutionary methods for optimizing both agents and the environments they learn in. A distinctive part of our agenda is designing methods that fit Apple's deployment reality — on-device and hybrid (device plus private cloud) execution, co-designed with current and future hardware — and that take advantage of what this ecosystem uniquely enables, such as deeply personalized, long-context agentic experiences. We aim for both field-changing research and direct impact on Apple products and internal engineering processes. MLR is a research group first. Management here is about spreading the load of running a team, not stepping away from the work — everyone, including leads, stays hands-on. We support continued engagement with the academic community: publishing, conference service, student collaboration, and internships.

Description

  • Lead research on RL and post-training for agentic capabilities: reward, preference optimization, and verifier design, training recipes, and evaluation for tool calling, coding, and multi-step interactive tasks.
  • Build and own synthetic data and task-generation pipelines — generating diverse, verifiable tasks and environments, along with the interactive environments and benchmarks that go with them, and the curricula that turn them into capable agents.
  • Drive codebases and infrastructure for the core RL research effort and help engage partner teams to use and co-develop the framework.
  • Manage and mentor a small team (3–4) of senior researchers and research engineers with distinct specialties, shaping a shared research direction while protecting room for bottom-up, idea-driven work.
  • Stay hands-on: run experiments, write code, and contribute directly to the team's most important technical problems.
  • Connect post-training research to efficiency and deployment: what works under on-device and hybrid compute constraints, and how method design interacts with hardware.
  • Collaborate across the organization on adjacent directions, including methods for environment and agent co-optimization, self-improvement, world models used as planners or policies, and personalized long-context agents.
  • Publish in top venues and engage with the broader research community.
Minimum Qualifications

PhD in machine learning or a related field, or equivalent research experience 7-10+ years of research experience beyond PhD in industry or as an academic research lead Strong track record in RL and/or post-training of large models, demonstrated through publications, open-source contributions, or shipped systems Leadership experience: setting and defending a research direction over multiple years, and directing others' work — through direct reports, PhD students, postdocs, or sustained project teams. Formal management experience is welcome but not required Experience owning ML infrastructure, frameworks and codebases, including open-source research frameworks or environment suites others build on Strong engineering skills; comfortable working hands-on in large training codebases

Preferred Qualifications

Experience taking research from idea to product or production impact Familiarity with efficiency-aware modeling: small models, mixture-of-experts, quantization, distillation, inference-cost constraints, or hardware-aware method design Interest or background in open-endedness, evolutionary computation, curriculum or environment design, multi-agent systems, or self-improving systems Principled or theoretical grounding in RL — representation, exploration, or optimization views of policy learning — alongside strong empirical work Breadth across core machine learning — generative models, self-supervised learning, pre-training — and perspective on the field's longer arcs, not only its most recent methods Experience growing other researchers, and managing researchers and engineers with heterogeneous specialties and synthesizing their work toward a common goal Experience owning a large RL or post-training codebase

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

AIML - Machine Learning Research Lead, RL Agents, MLR
AIML - Machine Learning Research Lead, RL Agents, MLR

Apple Inc. • Cupertino (CA), Northern (KY)

Hybrid
USD 216,000 - 394,000
Apple stock programs
Medical coverage
Retirement benefits
+1
Lead RL Researcher for Agentic AI & Environments
Lead RL Researcher for Agentic AI & Environments

Apple Inc. • Cupertino (CA), Northern (KY)

Hybrid
USD 216,000 - 394,000
Apple stock programs
Medical coverage
Retirement benefits
+1
Staff Software Engineer, Code RL
Staff Software Engineer, Code RL

Anthropic • New York (NY), Seattle (WA), San Francisco (CA)

On-site
USD 140,000 - 180,000
Member of Technical Staff, Research — Early Career(PHD)
Member of Technical Staff, Research — Early Career(PHD)

Goaly • Menlo Park (CA), Northern (KY)

Hybrid
USD 150,000 - 210,000
Meals and office benefits
Member of Technical Staff - Research & Post-training
Member of Technical Staff - Research & Post-training

Preference Model • San Francisco (CA)

On-site
USD 120,000 - 150,000
Competitive cash and equity compensation (>90th percentile)
Health, vision, dental benefits
401K match
+2
Member of Technical Staff, Post-training
Member of Technical Staff, Post-training

Hark • San Jose (CA)

On-site
USD 180,000 - 450,000
AIML - Machine Learning Enginner, Apple Foundation Models
AIML - Machine Learning Enginner, Apple Foundation Models

Socket.dev • Cary (NC)

On-site
USD 150,000 - 230,000
AIML - Machine Learning Engineer , Apple Foundation Models
AIML - Machine Learning Engineer , Apple Foundation Models

Apple • Cary (NC)

On-site
USD 190,000 - 240,000
RESEARCHER (GENERAL)
RESEARCHER (GENERAL)

MakerMaker.AI • San Francisco (CA)

On-site
USD 120,000 - 160,000
Research Engineer
Research Engineer

Bespoke Labs • Mountain View (CA)

On-site
USD 120,000 - 140,000
Health coverage
Opportunity to work with leading AI labs
Competitive salary and equity