AI Engineer – Algorithm Evaluation & Agentic Systems

Socket.dev

Sunnyvale (CA)

On-site

USD 140,000 - 190,000

Full time

10 days ago

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Apple DAQ team seeks an AI Engineer specializing in algorithm evaluation and agentic systems design for advanced computer vision and video understanding. You will lead benchmarking, develop evaluation frameworks, and build autonomous multi-modal workflows bridging experimentation with production.

The role requires a strong foundation in ML/CV, PyTorch, and Python, with experience in evaluating vision-language models and building scalable pipelines.

Qualifications

  • MS and at least 3 years of industry experience.
  • 3+ years applied ML/CV/AI system evaluation.
  • Strong ML foundations: probability, statistics, data distributions, and bias/variance.
  • CV expertise: ViTs, spatial-temporal modeling, image/video processing.

Responsibilities

  • Lead benchmarking and integration of state-of-the-art models for image and video understanding.
  • Design robust evaluation frameworks and metrics.
  • Develop autonomous multi-modal workflows bridging experimentation and production.
  • Collaborate to define metrics from first principles and communicate findings clearly.

Skills

Python
PyTorch
Computer Vision
ML Evaluation
Statistics
Edge-case analysis
Metrics design
System-level thinking
Communication

Education

MS in Computer Science / ML

Tools

Jupyter
Linux

Job description

How do we ensure Apple's next-generation AI products are robust, safe, and truly intelligent? Join the DAQ team to help answer that. We are seeking an AI Engineer specializing in algorithm evaluation and agentic systems design for advanced computer vision and video understanding algorithms.

What We Value
  • Production mindset: correctness, observability and maintainability
  • Ability to reason about system-level tradeoffs, not just model performance
  • Ability to balance experimentation speed with engineering rigor
  • Comfort working in ambiguous problem spaces and defining metrics from first principles
  • Clear communication of technical findings to both technical and non-technical audiences
Description

Within the DAQ team, our core mission is to evaluate and elevate advanced visual technologies. As a key member of this group, you will lead the benchmarking and integration of state-of-the-art models for image and video understanding. Rather than focusing on core model training, you will apply your deep CV and ML expertise to rigorously test models in applied settings, uncover edge-case failure modes, and architect advanced agentic systems. If you are passionate about AI safety, robust evaluation, and building autonomous multi-modal workflows that bridge experimentation with production, we’d love to hear from you.

Minimum Qualifications
  • MS and a minimum of 3 years relevant industry experience
  • 3+ years of applied experience in Machine Learning, Computer Vision, or AI System Evaluation
  • Solid ML Foundation: Deep understanding of core Machine Learning principles, including probability, statistics, data distributions, and model bias/variance. You can apply statistical rigor to ensure evaluation metrics are meaningful and reliable.
  • Computer Vision Expertise: Deep theoretical and practical understanding of Computer Vision (CV) and Vision-Language Models (VLMs). You must understand how Vision Transformers (ViTs), spatial-temporal modeling, and image/video processing work under the hood to effectively evaluate them.
  • Advanced Evaluation Skills: Proven track record of defining robust metrics/KPIs and designing rigorous evaluation frameworks for generative AI or foundation models. Deep experience with custom benchmark creation, automated regression testing, LLM/VLM-as-a-judge methodologies, and human-in-the-loop evaluation.
  • Agentic Systems: Experience building and evaluating LLM/VLM-powered agents, including tool use, multi-step reasoning, planning, and memory management workflows.
  • Failure Analysis: Strong intuition for probing ML models to discover edge cases, hallucinations, and performance bottlenecks in constrained environments. Be able to translate findings into actionable improvement recommendations.
  • Engineering Excellence: Strong proficiency in Python and experience with deep learning frameworks (PyTorch) for running inference, extracting embeddings, and building scalable evaluation pipelines.
Preferred Qualifications
  • Demonstrated ability to lead technical evaluation strategies end-to-end, drive architectural decisions for testing infrastructure, and mentor engineers.
  • Strong foundation in statistics, including hypothesis testing, confidence intervals, and experimental design
  • Knowledge of reinforcement learning, planning, or decision-making systems
  • Experience evaluating multi-modal or multi-agent systems
  • Prior work on AI reliability, safety, or benchmarking
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

AI Engineer - Algorithm Evaluation & Agentic Systems
AI Engineer - Algorithm Evaluation & Agentic Systems

Apple Inc. • Sunnyvale (CA)

On-site
USD 150,000 - 278,000
Stock programs
Education reimbursement
Comprehensive health benefits
AI Systems Evaluation Engineer - Vision & Agents
AI Systems Evaluation Engineer - Vision & Agents

Socket.dev • Sunnyvale (CA)

On-site
USD 140,000 - 190,000
AIML - Sr Engineering Specialist, Evaluation
AIML - Sr Engineering Specialist, Evaluation

Apple • Seattle (WA)

On-site
USD 130,000 - 180,000
Applied AI Engineer - iCloud Data
Applied AI Engineer - iCloud Data

Socket.dev • Cupertino (CA)

On-site
USD 210,000 - 320,000
Applied AI & Data Engineer - Business & Education
Applied AI & Data Engineer - Business & Education

Apple • Cupertino (CA)

On-site
USD 210,000 - 320,000
Agent Sciences and Infrastructure Engineer, AIML
Agent Sciences and Infrastructure Engineer, AIML

Socket.dev • Cupertino (CA)

On-site
USD 180,000 - 240,000
AIML - Machine Learning Engineer, Red Teaming - Responsible AI and Safety
AIML - Machine Learning Engineer, Red Teaming - Responsible AI and Safety

Apple • California (MO)

On-site
USD 180,000 - 230,000
Machine Learning Engineer – Computer Vision & Data Systems
Machine Learning Engineer – Computer Vision & Data Systems

Socket.dev • Seattle (WA)

On-site
USD 140,000 - 200,000
Senior Engineering Manager, Agentic AI & Intelligent Reasoning
Senior Engineering Manager, Agentic AI & Intelligent Reasoning

Socket.dev • Santa Clara (CA)

On-site
USD 180,000 - 350,000
AIML - Sr Applied AI Scientist - GenAI Model Autograding, Evaluation
AIML - Sr Applied AI Scientist - GenAI Model Autograding, Evaluation

Socket.dev • Cupertino (CA)

On-site
USD 180,000 - 240,000