Senior ML Evaluation Engineer for Agentic AI & Generative AI

Socket.dev

Cupertino (CA)

On-site

USD 180,000 - 240,000

Full time

3 days ago
Be an early applicant
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Job summary

Apple is seeking a Machine Learning Evaluation Engineer to design and scale evaluation systems for LLM, Generative AI, and agentic AI products. You will build datasets, tooling, and quality signals to measure accuracy, relevance, groundedness, and task completion across Commerce AI experiences.

You will collaborate with ML, software, product, QA, and data science teams to establish robust evaluation practices, develop Auto Eval capabilities, and integrate them into development pipelines for

Qualifications

  • Typically requires a minimum of 7 years of related experience in ML engineering, ML evaluation, software eng., data science, or a related field.
  • Experience designing automated ML evaluation frameworks, metrics, benchmarks, datasets, or experimentation methodologies.
  • Strong programming skills in Python and production-quality software development.
  • Understanding of modern LLM architectures, prompting, retrieval, and agentive workflows.
  • Experience with HITL evaluation, data-quality workflows, and statistical analysis.

Responsibilities

  • Design and build scalable evaluation systems for LLM and Generative AI products.
  • Define methodologies and quality metrics across accuracy, relevance, groundedness, and task completion.
  • Build and maintain evaluation datasets including golden datasets and regression suites.
  • Develop Auto Eval capabilities for evaluating models, prompts, and retrieval systems.
  • Design model-based evaluation approaches such as LLM-as-a-Judge and calibration methods.
  • Develop HITL evaluation approaches for subjective quality dimensions.
  • Create evaluation rubrics, annotation guidelines, and quality standards with cross-functional teams.
  • Calibrate automated evaluators against human judgment and measure reliability.
  • Evaluate end-to-end AI systems including context construction, prompts, and tool use.
  • Develop evaluation for multi-turn conversations, personalization, recommendations, and agentic tasks.
  • Perform detailed error analysis to identify model and data improvements.
  • Build reusable evaluation infrastructure, APIs, dashboards, and developer tooling.
  • Integrate evaluation into CI/CD workflows and release-readiness assessments.
  • Connect offline evaluation results with production signals to improve coverage.
  • Collaborate with ML, Software, Product, Quality, and Data Science teams throughout lifecycle.

Skills

Python
ML Engineering
Evaluation frameworks
Data pipelines
Production software
LLMs
Generative AI

Education

Bachelor's degree in CS/ML/AI
Master's degree preferred

Tools

CI/CD pipelines

Job description

Apple is seeking a Machine Learning Evaluation Engineer to design and scale evaluation systems for LLM, Generative AI, and agentic AI products. You will build datasets, tooling, and quality signals to measure accuracy, relevance, groundedness, and task completion across Commerce AI experiences.

You will collaborate with ML, software, product, QA, and data science teams to establish robust evaluation practices, develop Auto Eval capabilities, and integrate them into development pipelines for

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior ML Evaluation Engineer for Agentic AI Systems
Senior ML Evaluation Engineer for Agentic AI Systems

Apple Inc. • Cupertino (CA), Northern (KY)

Hybrid
USD 185,000 - 278,000
Medical and dental coverage
Retirement benefits
Employee stock purchase plan
+1
Machine Learning Engineer - Agentic AI Evaluation Frameworks
Machine Learning Engineer - Agentic AI Evaluation Frameworks

Socket.dev • Cupertino (CA)

On-site
USD 180,000 - 240,000
AI Evaluation Engineer for LLM & Multimodal Systems
AI Evaluation Engineer for LLM & Multimodal Systems

Apple Inc. • Cupertino (CA), Northern (KY)

Hybrid
USD 150,000 - 225,000
Medical and dental coverage
Employee stock programs
Stock purchase plan
+3
ML Engineer: AI Evaluation & LLM Systems
ML Engineer: AI Evaluation & LLM Systems

Socket.dev • Cupertino (CA)

On-site
USD 150,000 - 230,000
Machine Learning Engineer - Agentic AI Evaluation Frameworks
Machine Learning Engineer - Agentic AI Evaluation Frameworks

Apple Inc. • Cupertino (CA), Northern (KY)

Hybrid
USD 185,000 - 278,000
Medical and dental coverage
Retirement benefits
Employee stock purchase plan
+1
Human-Centered AI ML Engineer — Evaluation & Insights
Human-Centered AI ML Engineer — Evaluation & Insights

Socket.dev • Seattle (WA)

On-site
USD 150,000 - 190,000
AIML - Sr Applied AI Scientist - GenAI Model Autograding, Evaluation
AIML - Sr Applied AI Scientist - GenAI Model Autograding, Evaluation

Socket.dev • Cupertino (CA)

On-site
USD 180,000 - 240,000
AI Engineer — Agentic Systems & Vision Evaluation
AI Engineer — Agentic Systems & Vision Evaluation

Apple Inc. • Sunnyvale (CA)

On-site
USD 150,000 - 278,000
Stock programs
Education reimbursement
Comprehensive health benefits
AIML Evaluation Engineer - Shape Apple Intelligence
AIML Evaluation Engineer - Shape Apple Intelligence

Apple Inc. • Cupertino (CA)

On-site
USD 150,400 - 277,600
Medical and dental coverage
Retirement benefits
Employee stock programs
AIML - Sr Manager, Evaluation - Data Science & Insights
AIML - Sr Manager, Evaluation - Data Science & Insights

Socket.dev • Seattle (WA)

On-site
USD 180,000 - 240,000