AI Engineer - Algorithm Evaluation & Agentic Systems

Apple Inc.

Sunnyvale (CA)

On-site

USD 150,000 - 278,000

Full time

3 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Stock programs
Education reimbursement
Comprehensive health benefits

Job summary

Apple Inc. in Sunnyvale, California, seeks an AI Engineer specializing in algorithm evaluation and agentic systems for advanced computer vision and video understanding.

The role centers on rigorous testing, benchmarking, and building autonomous multi‑modal workflows bridging experimentation with production. You will lead evaluation pipelines, analyze failures, and collaborate with model training teams to steer iterations.

Qualifications

  • MS and a minimum of 3 years relevant industry experience.
  • 3+ years of applied experience in Machine Learning, Computer Vision, or AI System Evaluation.
  • Solid ML foundation: probability, statistics, data distributions, and model bias/variance.
  • Engineering with Python and PyTorch for running inference and building evaluation pipelines.

Responsibilities

  • Algorithm Evaluation & Benchmarking: design, build, and scale comprehensive evaluation pipelines.
  • Deep Failure Analysis: identify root causes of visual hallucinations, temporal inconsistencies, and edge-case failures.
  • Agentic Architecture: build, deploy, and evaluate agentic workflows that use vision models to solve multi-step problems.
  • Golden Data Curation: curate high-quality, schematized datasets and ground-truth benchmarks for multi-modal evaluation.
  • Cross‑Functional Collaboration: work with core model training teams to provide actionable metrics.

Skills

Machine Learning
Computer Vision
AI System Evaluation
Python
PyTorch
Evaluation Pipelines

Education

MS degree in a relevant field

Job description

AI Engineer – Algorithm Evaluation & Agentic Systems

How do we ensure Apple's next-generation AI products are robust, safe, and truly intelligent? Join the DAQ team to help answer that. We are seeking an AI Engineer specializing in algorithm evaluation and agentic systems design for advanced computer vision and video understanding algorithms.What We ValueProduction mindset: correctness, observability and maintainabilityAbility to reason about system-level tradeoffs, not just model performanceAbility to balance experimentation speed with engineering rigorComfort working in ambiguous problem spaces and defining metrics from first principlesClear communication of technical findings to both technical and non-technical audiences

Description

Within the DAQ team, our core mission is to evaluate and elevate advanced visual technologies. As a key member of this group, you will lead the benchmarking and integration of state-of-the‑art models for image and video understanding. Rather than focusing on core model training, you will apply your deep CV and ML expertise to rigorously test models in applied settings, uncover edge‑case failure modes, and architect advanced agentic systems. If you are passionate about AI safety, robust evaluation, and building autonomous multi‑modal workflows that bridge experimentation with production, we’d love to hear from you.

Responsibilities
  • Algorithm Evaluation & Benchmarking: Design, build, and scale comprehensive evaluation pipelines. You will be responsible for both holistic end‑to‑end system evaluation and granular component‑level testing to rigorously measure model capabilities on complex image and video understanding tasks.
  • Deep Failure Analysis: Leverage your CV and ML background to dive deep into model outputs, identifying root causes of visual hallucinations, temporal inconsistencies in video, and edge‑case failures.
  • Agentic Architecture: Build, deploy, and evaluate agentic workflows that utilize these vision models to autonomously solve multi‑step user problems (e.g., video summarization, visual search). You will heavily utilize component‑level evaluation to isolate and triage exactly which parts of the agentic workflow (e.g., tool selection, memory retrieval, visual reasoning) are succeeding or failing.
  • Golden Data Curation: Lead the strategy for curating high‑quality, schematized datasets and ground‑truth benchmarks specifically tailored for evaluating multi‑modal capabilities.
  • Cross‑Functional Collaboration: Partner closely with the core model training teams. You will provide them with actionable, data‑driven insights and metrics to guide the next iteration of model training and fine‑tuning.
Minimum Qualifications
  • MS and a minimum of 3 years relevant industry experience
  • 3+ years of applied experience in Machine Learning, Computer Vision, or AI System Evaluation
  • Solid ML Foundation: Deep understanding of core Machine Learning principles, including probability, statistics, data distributions, and model bias/variance. You can apply statistical rigor to ensure evaluation metrics are meaningful and reliable.
  • Computer Vision Expertise: Deep theoretical and practical understanding of Computer Vision (CV) and Vision‑Language Models (VLMs). You must understand how Vision Transformers (ViTs), spatial‑temporal modeling, and image/video processing work under the hood to effectively evaluate them.
  • Advanced Evaluation Skills: Proven track record of defining robust metrics/KPIs and designing rigorous evaluation frameworks for generative AI or foundation models. Deep experience with custom benchmark creation, automated regression testing, LLM/VLM‑as‑a‑judge methodologies, and human‑in‑the‑loop evaluation.
  • Agentic Systems: Experience building and evaluating LLM/VLM‑powered agents, including tool use, multi‑step reasoning, planning, and memory management workflows.
  • Failure Analysis: Strong intuition for probing ML models to discover edge cases, hallucinations, and performance bottlenecks in constrained environments. Be able to translate findings into actionable improvement recommendations.
  • Engineering Excellence: Strong proficiency in Python and experience with deep learning frameworks (PyTorch) for running inference, extracting embeddings, and building scalable evaluation pipelines.
Preferred Qualifications
  • Demonstrated ability to lead technical evaluation strategies end‑to‑end, drive architectural decisions for testing infrastructure, and mentor engineers.
  • Strong foundation in statistics, including hypothesis testing, confidence intervals, and experimental design
  • Knowledge of reinforcement learning, planning, or decision‑making systems
  • Experience evaluating multi‑modal or multi‑agent systems
  • Prior work on AI reliability, safety, or benchmarking

At Apple, base pay is one part of our total compensation package and is determined within a range. This provides the opportunity to progress as you grow and develop within a role. The base pay range for this role is between $150,400 and $277,600, and your base pay will depend on your skills, qualifications, experience, and location.

Apple employees also have the opportunity to become an Apple shareholder through participation in Apple’s discretionary employee stock programs. Apple employees are eligible for discretionary restricted stock unit awards, and can purchase Apple stock at a discount if voluntarily participating in Apple’s Employee Stock Purchase Plan. You’ll also receive benefits including: Comprehensive medical and dental coverage, retirement benefits, a range of discounted products and free services, and for formal education related to advancing your career at Apple, reimbursement for certain educational expenses — including tuition. Additionally, this role might be eligible for discretionary bonuses or commission payments as well as relocation. Learn more about Apple Benefits

Note: Apple benefit, compensation and employee stock programs are subject to eligibility requirements and other terms of the applicable plan or program.

Apple is an equal opportunity employer that is committed to inclusion and diversity. We seek to promote equal opportunity for all applicants without regard to race, color, religion, sex, sexual orientation, gender identity, national origin, disability, Veteran status, or other legally protected characteristics. Learn more about your EEO rights as an applicant

At Apple, we believe accessibility is a fundamental human right. You’ll find that idea reflected in everything here — in our culture, our benefits and our digital tools. By welcoming as many perspectives as possible, we help you build a career where you feel like you belong.

Learn about accessibility in Apple’s workplace

Learn about reasonable accommodations for job applicants

Apple accepts applications to this posting on an ongoing basis.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

AI Engineer – Algorithm Evaluation & Agentic Systems
AI Engineer – Algorithm Evaluation & Agentic Systems

Socket.dev • Sunnyvale (CA)

On-site
USD 140,000 - 190,000
AI/ML Software Engineer
AI/ML Software Engineer

Apple Inc. • Sunnyvale (CA)

On-site
USD 150,000 - 278,000
Medical and dental coverage
Stock programs
Relocation support
Evaluation & Insights Machine Learning Engineer
Evaluation & Insights Machine Learning Engineer

Apple Inc. • Cupertino (CA)

On-site
USD 184,000 - 325,000
AIML - AI Software Engineer, Evaluation
AIML - AI Software Engineer, Evaluation

Apple Inc. • Cupertino (CA)

On-site
USD 147,000 - 273,000
Comprehensive medical and dental coverage
Retirement benefits
Employee stock programs
+1
Applied AI Engineer - iCloud Data
Applied AI Engineer - iCloud Data

Apple Inc. • Seattle (WA), Northern (KY)

Hybrid
USD 175,000 - 309,000
AIML - Sr Machine Learning Engineer, Evaluation
AIML - Sr Machine Learning Engineer, Evaluation

Apple Inc. • Cupertino (CA)

On-site
USD 212,000 - 387,000
Medical and dental coverage
Retirement benefits
Employee stock programs
+2
Applied AI Engineer
Applied AI Engineer

Apple Inc. • Cupertino (CA)

On-site
USD 184,000 - 278,000
Applied AI & Data Engineer - Business & Education
Applied AI & Data Engineer - Business & Education

Apple Inc. • Cupertino (CA), Northern (KY)

Hybrid
USD 185,000 - 325,000
Apple Benefits
Relocation assistance
Discretionary bonuses
AIML - Sr Manager, Evaluation - Data Science & Insights
AIML - Sr Manager, Evaluation - Data Science & Insights

Apple Inc. • Seattle (WA), Northern (KY)

Hybrid
USD 226,000 - 382,000
AIML - Data Scientist, Responsible AI, Product Insights
AIML - Data Scientist, Responsible AI, Product Insights

Apple Inc. • Cupertino (CA)

On-site
USD 181,000 - 319,000
Comprehensive medical and dental coverage
Employee Stock Purchase Plan
Tuition reimbursement for formal education
+1