AI Evaluations Engineer, US Decision Intelligence

Apple Inc.

Cupertino (CA)

On-site

USD 185,000 - 278,000

Full time

5 days ago
Be an early applicant
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Benefits offered by this job

Medical insurance
Dental coverage
Retirement benefits
Discounted products
Free services
Tuition reimbursement

Job summary

Apple Inc. in Cupertino, CA is seeking an AI Evaluations Engineer to own the end-to-end evaluation pipeline for AI products and agentic workflows. You will build and operate evaluation frameworks, instrumentation, and rubrics to measure accuracy, relevance, and latency across LLM outputs and agent actions.

The role requires deep experience with AI evaluation techniques, multiple LLM ecosystems, and collaboration with cross-functional teams to drive product quality and reliability.

Qualifications

  • 5+ years in data and AI-related fields such as AI engineering or data science.
  • Experience with AI evaluation techniques like rubric-based scoring.
  • Familiarity with multiple LLM ecosystems (OpenAI, Anthropic, Gemini).
  • SQL proficiency and data analytics platform exposure (Hadoop/Spark/Snowflake).
  • CI/CD or release validation experience and cross-team collaboration.

Responsibilities

  • Architect end-to-end AI evaluation frameworks and AB testing strategies.
  • Build and operate evaluation workflows for LLM outputs across tasks.
  • Define rubric-based scoring for correctness, relevance, grounding, and consistency.
  • Instrument traces, prompts, responses, and user feedback for observability.
  • Own the eval gate and pass/fail contracts for releases.
  • Collaborate with AI engineers and platform teams to map requirements to criteria.

Skills

AI evaluation
SQL
CI/CD
Langfuse
Data analysis
Cross-functional

Education

B.S. in CS/Engineering
MS/PhD preferred

Tools

Pinecone
FAISS
Milvus
PostgreSQL
RabbitMQ
Redis
Valkey
Langfuse

Job description

AI Evaluations Engineer, US Decision Intelligence

Cupertino, California, United States Machine Learning and AI

Imagine what you could do here. At Apple, new ideas have a way of becoming outstanding products, services, and customer experiences very quickly. Bring passion and dedication to your job, and there's no telling what you could accomplish.Apple’s Sales organization generates the revenue needed to fuel our ongoing development of products and services. This, in turn, enriches the lives of hundreds of millions of people around the world. We are, in many ways, the face of Apple to our largest customers.Apple's US Decision Intelligence (DI) team is looking for a talented individual who is passionate about crafting, implementing, and operating AI solutions that have a direct and measurable impact on Apple Sales and its customers.

Description

We’re seeking a visionary AI Evaluations Engineer to own the end-to-end evaluation pipeline for our AI products and agentic workflows. This role will focus on implementing and maintaining evaluation frameworks, instrumentation, and workflows that help us understand how well our AI systems perform, where they fail, and how they improve over time. You own the evaluation gate and the standards.This role will operate in both capacities, to augment existing AI roadmap, as well as innovate and trailblaze new frontier-technology projects, crafting AI experiences that reduce time to insight and catalyze decision making.

Responsibilities
  • Architect a comprehensive Evals framework to trace at every layer, from agent responses breakdown, to skill level, and tool calling with the main objective to optimize for accuracy and performance improving latency and running AB tests to find the best recommendation across the system.
  • Build and operate AI evaluation workflows that measure the quality of LLM outputs across chat, summarization, recommendations, and agentic actions.
  • Implement rubric-based evals to score outputs for correctness, relevance, grounding, and consistency.
  • Move beyond LLM-as-a-judge to agent-as-a-judge, including harness-as-a-judge patterns applied against real traces in a sandbox.
  • Instrument LLM and agent workflows to capture traces, prompts and responses, metadata, and user feedback.
  • Own the platform-wide eval gate: define the pass/fail contract every track ships against, and hold a release when it isn't met.
  • Help define agent-specific evaluations (task completion, tool correctness, error recovery).
  • Partner with AI engineers and AI platform teams to translate product requirements into evaluation criteria.
  • Arbitrate eval disputes with track leads and own the standard that geo eval engineers implement agains.
  • Define rerun policy and variance thresholds.
  • Collaborate with the Evals team in India to maximize global impact.
  • Contribute to system design for observability, retries, and logging.
Minimum Qualifications
  • 5+ years of experience in data and AI-related fields such as AI engineering, software development, ML engineering, data science, or QA roles.
  • Eagerness and ability to learn new skills and solve dynamic problems in an encouraging and expansive environment.
  • Hands-on experience with AI evaluation techniques, such as Golden datasets, LLM-as-a-Judge, or rubric-based scoring.
  • Experience with different LLM ecosystems (OpenAI, Anthropic, Gemini, etc.), RAG pipelines, vector databases (e.g., Pinecone, FAISS, Milvus, PostgreSQL).
  • Proficiency in SQL and experience with at least one major data analytics platform, such as Hadoop, Spark, or Snowflake.
  • Experience with CI/CD or release validation workflows.
  • Experience working with data science teams on insights generation leveraging LLMs.
  • Strong time management skills with the ability to collaborate across multiple teams.
  • Able to balance competing priorities, long-term projects, and ad hoc requirements.
  • Ability to work in a fast-paced, dynamic, constantly evolving business environment.
  • Hands-on experience with Langfuse or similar tools for LLM observability.
  • Comfortable working with product/domain experts to translate fuzzy correctness criteria into measurable rubrics or metrics.
  • B.S. degree in Computer Science/Engineering, or equivalent work experience
Preferred Qualifications
  • Sound communication skills - expert at messaging domain and technical content, at a level appropriate for the audience. Strong ability to gain trust with stakeholders and senior leadership.
  • Familiarity with embeddings, retrieval algorithms, agents, and data modeling for vector and graph databases.
  • Other complementary technologies for distributed systems architecture and asynchronous messaging, agent communication, and caching like RabbitMQ, Redis, and Valkey are preferred.
  • Experience working across global teams to ensure alignment of product development.
  • Applied knowledge of GenAI and RAG strategies, microservices, recommendation systems, and context engineering.
  • Working knowledge of agent evaluation concepts like trajectory vs. end-to-end vs. component-level evaluation, tool-call correctness.
  • Advanced degree (MS or Ph.D.) in Economics, Electrical Engineering, Statistics, Data Science, or a similar quantitative field is preferred.

At Apple, base pay is one part of our total compensation package and is determined within a range. This provides the opportunity to progress as you grow and develop within a role. The base pay range for this role is between $184,700 and $277,600, and your base pay will depend on your skills, qualifications, experience, and location.

Apple employees also have the opportunity to become an Apple shareholder through participation in Apple’s discretionary employee stock programs. Apple employees are eligible for discretionary restricted stock unit awards, and can purchase Apple stock at a discount if voluntarily participating in Apple’s Employee Stock Purchase Plan. You’ll also receive benefits including: Comprehensive medical and dental coverage, retirement benefits, a range of discounted products and free services, and for formal education related to advancing your career at Apple, reimbursement for certain educational expenses — including tuition. Additionally, this role might be eligible for discretionary bonuses or commission payments as well as relocation. Learn more about Apple Benefits

Note: Apple benefit, compensation and employee stock programs are subject to eligibility requirements and other terms of the applicable plan or program.

Apple is an equal opportunity employer that is committed to inclusion and diversity. We seek to promote equal opportunity for all applicants without regard to race, color, religion, sex, sexual orientation, gender identity, national origin, disability, Veteran status, or other legally protected characteristics. Learn more about your EEO rights as an applicant

At Apple, we believe accessibility is a fundamental human right. You’ll find that idea reflected in everything here — in our culture, our benefits and our digital tools. By welcoming as many perspectives as possible, we help you build a career where you feel like you belong.

Learn about accessibility in Apple’s workplace

Learn about reasonable accommodations for job applicants

Apple accepts applications to this posting on an ongoing basis.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

AIML - Software Engineer - AI, Evaluation
AIML - Software Engineer - AI, Evaluation

Apple Inc. • Cupertino (CA), Northern (KY)

Hybrid
USD 150,000 - 278,000
Comprehensive medical and dental cover
Retirement benefits
Employee stock purchase plan
+1
AIML - Director, Evaluation - Data Science and Insights
AIML - Director, Evaluation - Data Science and Insights

Apple Inc. • Cupertino (CA), Northern (KY)

Hybrid
USD 311,000 - 497,000
Medical and dental coverage
Retirement benefits
Stock purchase plan
+2
AIML - Sr Manager, Evaluation - Data Science & Insights
AIML - Sr Manager, Evaluation - Data Science & Insights

Apple Inc. • Seattle (WA), Northern (KY)

On-site
USD 226,000 - 382,000
Machine Learning Engineer, Human Centered AI - Evaluations & Insights
Machine Learning Engineer, Human Centered AI - Evaluations & Insights

Apple Inc. • Seattle (WA)

On-site
USD 142,000 - 263,000
Sr. Applied Scientist, AI Evaluation & Quality Systems
Sr. Applied Scientist, AI Evaluation & Quality Systems

Apple Inc. • Seattle (WA)

On-site
USD 142,000 - 263,000
Apple Benefits
Relocation assistance
Employee stock purchase plan
+2
AIML - Sr Data Scientist, Evaluation
AIML - Sr Data Scientist, Evaluation

Apple Inc. • Seattle (WA)

On-site
USD 184,700 - 324,800
Stock-based compensation
Medical and dental coverage
Relocation assistance
AIML - Sr Machine Learning Engineering Manager, Evaluation
AIML - Sr Machine Learning Engineering Manager, Evaluation

Apple Inc. • Cupertino (CA)

On-site
USD 238,000 - 402,000
Stock options
Relocation assistance
Comprehensive medical/dental coverage
+1
AIML - Data Scientist, Evaluation
AIML - Data Scientist, Evaluation

Apple • Cupertino (CA)

On-site
USD 147,000 - 273,000
Employee stock programs
Comprehensive medical coverage
Retirement benefits
+1
AI Engineer – Algorithm Evaluation & Agentic Systems
AI Engineer – Algorithm Evaluation & Agentic Systems

Apple Inc. • Sunnyvale (CA)

Hybrid
USD 150,000 - 278,000
Apple benefits
Applied AI & Data Engineer - Business & Education
Applied AI & Data Engineer - Business & Education

Apple Inc. • Cupertino (CA), Northern (KY)

Hybrid
USD 185,000 - 325,000
Apple Benefits
Relocation assistance
Discretionary bonuses