AIML - Sr Machine Learning Engineer, Evaluation

Apple Inc.

Cupertino (CA)

On-site

USD 212,000 - 386,300

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Medical and dental coverage
Retirement benefits
Employee stock programs
Discounted products and services
Educational reimbursement

Job summary

Apple Inc. is seeking a Senior Machine Learning Engineer in Cupertino, California, to evaluate and refine Apple's AI systems. You will design and develop key infrastructures for model and agent evaluations, contribute to quality improvements, and work closely with product teams.

The role requires at least 8 years of experience in software engineering, a strong background in machine learning, and proficiency in Python and ML frameworks such as PyTorch. Join Apple to advance AI products impacting millions of users.

Qualifications

  • 8+ years of professional experience as a software engineer.
  • Strong background in machine learning and distributed systems.
  • Experience with ML infrastructure for evaluation, training, or deployment.

Responsibilities

  • Design and build evaluation infrastructure for agents and foundation models.
  • Develop LLM judges and reward models.
  • Collaborate with product teams on quality gaps.

Skills

Machine learning systems
Distributed infrastructure
Problem-solving skills
Collaboration
Python
LLM evaluation

Education

Bachelor's or Master's degree in Computer Science

Tools

PyTorch

Job description

AIML - Sr Machine Learning Engineer, Evaluation

Cupertino, California, United States Machine Learning and AI

We are seeking a highly skilled and experienced machine learning engineer to join AIML Evaluation to build the systems that evaluate and refine Apple's foundation models and agents. As a key member of the team, you will help design and develop benchmarks, evaluators, simulation environments, and prompt and context optimization pipelines that drive quality improvements across Apple's AI experiences. You will collaborate with product teams and the foundation model team to close the loop between observation and improvement, contributing datasets, environments, and reward signals that drive model and agent quality.

Description

Our team builds the benchmarks, environments, and tooling that power model and agent refinement, and turns observations into actionable opportunities for the next model and agent iteration. We work across the full spectrum of evaluation: offline benchmarks, device-in-the-loop simulation, and on-device observation in production. We develop LLM-as-judge evaluators, train reward models calibrated against human feedback, optimize prompts and context for agents, and contribute targeted datasets and reward signals to foundation model post‑training.

In this role, you will play a crucial role in designing and developing evaluation and refinement infrastructure that supports a broad range of AI products at Apple. You will work on agent and model evaluation across offline, device-in-the-loop, and on-device settings; build automated prompt and context optimization pipelines; and partner with product and research teams to translate failure analysis into measurable model and agent improvements. You will also have the opportunity to engage with product teams across Apple and contribute to advancements in large language models and agentic systems that will reach millions of users.

To succeed in this role, you should have a strong background in machine learning systems, distributed infrastructure, and a proven track record of building and maintaining ML evaluation or training infrastructure. You should be a proactive problem solver with excellent communication skills and the ability to work effectively across multiple codebases, teams, and organizations. Experience with LLM evaluation, reward modeling, prompt optimization, or agentic systems is highly desirable.

Responsibilities
  • Design and build evaluation infrastructure for agents and foundation models.
  • Develop LLM judges, reward models, and prompt optimization pipelines.
  • Build and integrate simulation environments for agent evaluation and trajectory-based data generation.
  • Collaborate with product teams to identify, prioritize, and address quality gaps.
  • Contribute datasets, environments, and reward signals to the foundation model post‑training loop.
Minimum Qualifications
  • Strong background in machine learning and distributed systems.
  • Experience building and maintaining ML infrastructure for evaluation, training, or deployment.
  • Ability to work effectively across multiple codebases, teams, and organizations.
  • 8+ years of professional experience as a software engineer, preferably in machine learning or a related field.
  • Bachelor's or Master's degree in Computer Science or a related field.
Preferred Qualifications
  • Experience with LLM evaluation, LLM-as-judge, or reward modeling.
  • Experience with prompt optimization, agent harness development, or post-training (SFT, DPO, RLHF).
  • Proficiency in Python and ML frameworks such as PyTorch.
  • Experience with agentic systems, simulation environments, or trajectory-based data generation.
  • Familiarity with on-device or privacy-preserving ML.
  • Proactive and determined problem-solving skills.

At Apple, base pay is one part of our total compensation package and is determined within a range. The base pay range for this role is between $212,000 and $386,300, and your base pay will depend on your skills, qualifications, experience, and location.

Apple employees also have the opportunity to become shareholders through participation in Apple’s discretionary employee stock programs. Apple employees are eligible for discretionary restricted stock unit awards and can purchase Apple stock at a discount if voluntarily participating in Apple’s Employee Stock Purchase Plan. You’ll also receive benefits including comprehensive medical and dental coverage, retirement benefits, a range of discounted products and free services, and reimbursement for certain educational expenses— including tuition. Additionally, this role might be eligible for discretionary bonuses or commission payments as well as relocation.

Note: Apple benefit, compensation and employee stock programs are subject to eligibility requirements and other terms of the applicable plan or program.

Apple is an equal opportunity employer that is committed to inclusion and diversity. We seek to promote equal opportunity for all applicants without regard to race, color, religion, sex, sexual orientation, gender identity, national origin, disability, Veteran status, or other legally protected characteristics.

We believe accessibility is a fundamental human right. By welcoming as many perspectives as possible, we help you build a career where you feel like you belong.

Apple accepts applications to this posting on an ongoing basis.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

AIML - AI Software Engineer, Evaluation
AIML - AI Software Engineer, Evaluation

Apple Inc. • Cupertino (CA)

On-site
USD 147,000 - 273,000
Comprehensive medical and dental coverage
Retirement benefits
Employee stock programs
+1
AIML - Sr Manager, Evaluation - Data Science & Insights
AIML - Sr Manager, Evaluation - Data Science & Insights

Apple Inc. • Seattle (WA), Northern (KY)

Hybrid
USD 226,000 - 382,000
AIML - Machine Learning Engineer, Foundation Models
AIML - Machine Learning Engineer, Foundation Models

Apple Inc. • Cupertino (CA)

On-site
USD 181,000 - 319,000
Comprehensive medical and dental coverage
Retirement benefits
Employee stock purchase program
+1
AIML - Sr Machine Learning Engineer, Data and ML Innovation
AIML - Sr Machine Learning Engineer, Data and ML Innovation

Apple Inc. • Cupertino (CA)

On-site
USD 150,000 - 278,000
Employee stock programs
Discretionary bonuses
Relocation assistance
+2
AIML - Director, Data Science and Insights
AIML - Director, Data Science and Insights

Apple Inc. • Cupertino (CA)

On-site
USD 305,000 - 488,000
Comprehensive medical and dental coverage
Retirement benefits
Employee stock purchase plan
+1
AIML - Machine Learning Engineer, Foundation Models
AIML - Machine Learning Engineer, Foundation Models

Apple Inc. • Seattle (WA)

On-site
USD 184,700 - 324,800
Medical and dental coverage
Employee stock programs
Relocation assistance
Evaluation & Insights Machine Learning Engineer
Evaluation & Insights Machine Learning Engineer

Apple Inc. • Cupertino (CA)

On-site
USD 184,000 - 325,000
AIML - Machine Learning Researcher, Data and ML Innovation
AIML - Machine Learning Researcher, Data and ML Innovation

Apple Inc. • Santa Clara (CA)

On-site
USD 181,000 - 319,000
Comprehensive medical and dental coverage
Retirement benefits
Discounted products and free services
+1
AIML - Data Scientist, Evaluation
AIML - Data Scientist, Evaluation

Apple • Cupertino (CA)

On-site
USD 147,000 - 273,000
Employee stock programs
Comprehensive medical coverage
Retirement benefits
+1
AIML - Sr Engineering Specialist, Evaluation
AIML - Sr Engineering Specialist, Evaluation

Apple Inc. • Seattle (WA), Northern (KY)

Hybrid
USD 115,000 - 236,000
Stock programs
Restricted Stock Units (RSU)
Medical and dental coverage
+4