Machine Learning Engineer, ML/GenAI Evaluation

Apple Inc.

Austin, Northern (TX, KY)

Hybrid

USD 140,000 - 190,000

Full time

31 hours ago
Be an early applicant
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Job summary

Apple Inc. seeks a Machine Learning Engineer focused on Evaluation for Wallet, Payments, and Commerce. You will define the evaluation framework, bias testing, and readiness gates ensuring high accuracy, robustness, and reliability for hundreds of millions of users across global markets.

The role emphasizes adversarial testing, audit-like sign-off processes, and collaboration with ML, Product, Privacy, and Legal teams to raise the quality bar for model deployment.

Qualifications

  • MS in ML or related field preferred.
  • 7+ years in ML evaluation or related work.
  • 5+ years ML experience focused on evaluation.
  • Experience designing evaluation frameworks beyond accuracy.
  • Strong Python and experiment-tracking proficiency.

Responsibilities

  • Define evaluation criteria and quality metrics for Wallet features.
  • Design test sets covering diverse real-world scenarios and adversarial inputs.
  • Develop robustness testing methodologies for distribution shift and drift.
  • Own fairness evaluation end-to-end and gate launches on fairness criteria.
  • Build benchmarks reflecting Wallet's global user base and spending patterns.
  • Evaluate generative model outputs and establish sign-off criteria for launches.
  • Synthesize results into actionable insights for product decisions.
  • Collaborate with ML and Quality engineers to address failure modes early.

Skills

ML evaluation
Data analysis
Communication
Cross-functional collaboration

Education

MS in Machine Learning
PhD (preferred)

Tools

Python
MLflow
Weights & Biases
OCR pipelines

Job description

Machine Learning Engineer, ML/GenAI Evaluation

Austin, Texas, United States Software and Services

Would you like to contribute to Machine Learning and Generative AI technologies? Are you passionate about measuring what matters and ensuring AI systems work reliably for everyone? Do you believe that rigorous evaluation — including holding models accountable to fairness standards — is what separates great ML from good ML? We truly believe it is! We are defining what exceptional looks like for machine learning across Wallet, Payments, and Commerce. As a Machine Learning Engineer specializing in Evaluation, you will establish the evaluation criteria, metrics frameworks, and quality standards that determine when models are ready to reach hundreds of millions of users. Your judgment shapes model quality and earns the confidence to ship. You'll work at the intersection of rigorous ML science and high-impact product decisions, collaborating closely with ML Engineering, Product, Privacy, and Legal teams. This unique opportunity puts you at the center of model quality — designing adversarial test strategies, surfacing failure modes before they reach users, and owning the sign-off process that ensures Apple's financial features meet the highest bar for accuracy, robustness, and reliability.

Description

The ideal candidate is a rigorous, curious ML practitioner who believes that how you measure a model is just as important as how you train it. You think critically about what metrics actually capture, know how models break in the real world, and hold quality standards others find uncomfortably high — including on dimensions like fairness. You will own the full evaluation lifecycle for ML models across Wallet features — designing test frameworks, adversarial corpora, and benchmarks that reflect the diversity of Apple's global user base, then making the final quality call before any model ships. Your findings directly shape model development priorities and product decisions at scale.

Responsibilities
  • Define evaluation criteria and quality metrics for ML models powering Wallet features
  • Design and maintain structured test sets covering the full diversity of real-world scenarios — varied document formats, distributions, languages, edge cases, and adversarial inputs.
  • Develop evaluation methodologies for robustness testing: distribution shift, out-of-distribution generalization, temporal drift, and aggressor scenarios
  • Own fairness evaluation end-to-end — define fairness metrics appropriate to each Wallet feature, build bias test suites across protected attributes and user populations, measure disparate performance across subgroups, and gate model launches on fairness criteria with the same rigor as other conventional metrics.
  • Build user persona–stratified benchmarks that reflect the breadth of Wallet's global user population across spending patterns, locales, and document types
  • Evaluate generative and agentic model outputs — assessing hallucination rates, faithfulness, and groundedness using LLM-as-a-judge frameworks, human evaluation protocols, and prompt regression testing
  • Own model quality sign-off — establish the launch criteria, run final evaluations, and make the call on model readiness before any feature ships
  • Synthesize evaluation results into clear, actionable insights that guide model development priorities and product decisions
  • Partner with ML engineers and Quality engineers to identify failure modes early in the development cycle and close the loop between evaluation findings and model improvements
  • Establish and evangelize evaluation best practices across the Wallet ML team, raising the quality bar for how models are tested, monitored, and maintained post-launch
Minimum Qualifications
  • M.S. in Machine Learning, Computer Science, Statistics, Applied Mathematics, or a related technical field strongly preferred.
  • Bachelor's degree with 7+ years hands-on experience in ML evaluation, model quality, or applied research will be considered
  • 5+ years of hands-on ML experience, with deep expertise in model evaluation, offline metrics design, and behavioral testing
  • Strong track record designing evaluation frameworks for production ML systems — not just accuracy/F1, but precision-recall tradeoffs, calibration, fairness, and task-specific quality dimensions
  • Creative mindset with the ability to translate standard ML evaluation metrics (F1, AUC, etc.) into utility and user trust measures
  • Experience testing for distribution shift, out-of-distribution generalization, and temporal drift in real-world deployed models
  • Proven ability to construct adversarial test suites, aggressor scenarios, and edge-case corpora that surface model failure modes before they reach users
  • Experience with structured and semi-structured document understanding, OCR pipelines, or financial data extraction is a strong plus
  • Strong programming skills in Python; fluency with evaluation tooling, data pipelines, and experiment tracking (e.g., MLflow, W&B, or equivalent)
  • Excellent communication skills — ability to translate metric results into product-quality narratives for engineering and executive audiences
  • Experience owning model quality sign-off in a cross-functional launch process
Preferred Qualifications
  • PhD in Computer Science, Data Science, Statistics, AI/ML, or a related field.
  • Experience with Bayesian or causal graph-based approaches to data generation.
  • Experience with causal approaches to fairness evaluation — counterfactual fairness, causal Shapley values, or structural causal model–based bias auditing.
  • Experience evaluating models under privacy constraints or on-device inference settings is a plus.
  • Familiarity with confidence calibration techniques and uncertainty quantification a plus
  • Background in financial services, fintech, or consumer payment products

Apple is an equal opportunity employer that is committed to inclusion and diversity. We seek to promote equal opportunity for all applicants without regard to race, color, religion, sex, sexual orientation, gender identity, national origin, disability, Veteran status, or other legally protected characteristics. Learn more about your EEO rights as an applicant

At Apple, we believe accessibility is a fundamental human right. You’ll find that idea reflected in everything here — in our culture, our benefits and our digital tools. By welcoming as many perspectives as possible, we help you build a career where you feel like you belong.
Learn about accessibility in Apple’s workplace
Learn about reasonable accommodations for job applicants

Apple accepts applications to this posting on an ongoing basis.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Machine Learning Engineer, ML/GenAI Evaluation
Machine Learning Engineer, ML/GenAI Evaluation

Apple Inc. • San Diego (CA)

On-site
USD 175,000 - 310,000
Apple Benefits
Stock options
Tuition reimbursement
+2
Data Scientist, AI/ML Model Quality
Data Scientist, AI/ML Model Quality

Apple Inc. • New York (NY)

On-site
USD 130,000 - 170,000
Machine Learning Engineer, Human Centered AI - Evaluations & Insights
Machine Learning Engineer, Human Centered AI - Evaluations & Insights

Apple Inc. • Seattle (WA)

On-site
USD 142,000 - 263,000
Sr. Applied Scientist, AI Evaluation & Quality Systems
Sr. Applied Scientist, AI Evaluation & Quality Systems

Apple Inc. • Seattle (WA)

On-site
USD 142,000 - 263,000
Apple Benefits
Relocation assistance
Employee stock purchase plan
+2
Machine Learning Engineer - AI Evaluation & LLM Systems
Machine Learning Engineer - AI Evaluation & LLM Systems

Apple Inc. • Cupertino (CA), Northern (KY)

On-site
USD 150,000 - 225,000
Medical and dental coverage
Employee stock programs
Stock purchase plan
+3
AIML - Sr Software Engineer - AI, Evaluation
AIML - Sr Software Engineer - AI, Evaluation

Apple Inc. • Cupertino (CA), Northern (KY)

Hybrid
USD 150,000 - 278,000
Relocation
Apple stock programs
Medical and dental coverage
+1
AIML - Sr Machine Learning Engineering Manager, Evaluation
AIML - Sr Machine Learning Engineering Manager, Evaluation

Apple Inc. • Cupertino (CA)

On-site
USD 238,000 - 402,000
Stock options
Relocation assistance
Comprehensive medical/dental coverage
+1
AIML - Sr Machine Learning Engineer, DMLI
AIML - Sr Machine Learning Engineer, DMLI

Apple Inc. • Cupertino (CA)

On-site
USD 150,000 - 278,000
Medical and dental coverage
Retirement benefits
Employee stock programs
+2
Senior Machine Learning Engineer
Senior Machine Learning Engineer

Apple Inc. • Seattle (WA), Northern (KY)

On-site
USD 175,000 - 309,000
Machine Learning Engineer, Wallet Intelligence and Machine Learning
Machine Learning Engineer, Wallet Intelligence and Machine Learning

Apple Inc. • Cupertino (CA)

On-site
USD 167,000 - 278,000
Medical and dental coverage
Retirement benefits
Discounted Apple products
+1