Senior Software Development Engineer in Test — LLM Evaluation & Automation, T3E

Apple

San Diego (CA)

On-site

USD 140,000 - 190,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Apple seeks a Senior SDET to lead automated model evaluation across our generative AI pipelines, standing up LLM-as-a-judge and integrating it into CI/CD workflows.

You will build end-to-end eval coverage, curate rubrics, run scalable eval jobs, and help surface reliable quality signals before human review or deployment.

This hands-on, senior individual-contributor role requires strong Python, testing, and collaboration with modeling, framework, and infra teams.

Qualifications

  • BS in Computer Science, Mathematics, or a related field (or equivalent practical experience)
  • Three years of relevant industry experience in test automation, software development, or related areas.

Responsibilities

  • Design, build, and maintain LLM-as-a-judge evaluation harnesses and integrate them into existing and new CI/automation pipelines.
  • Author and curate eval sets and rubrics; partner with modeling teams whose own eval sets can run hundreds of examples judged by a separate model.
  • Run eval jobs at scale, triage results, and distinguish real model regressions from rubric problems or infrastructure noise so the signal stays actionable.
  • Plan and build test coverage and tooling from low level component tests through to end-to-end tests that exercise models and frameworks powering Generative AI features.
  • Package eval tooling for reuse — reusable libraries and jobs that other engineers on the team and partner teams can adopt.
  • Define quality gates and reporting so model regressions are caught and communicated before human eval or population rollout.
  • Collaborate with data scientists and modeling engineers on approaches to spot regressions across large output sets or between model updates.

Skills

Python
CI/CD
Test automation
Debugging
Communication

Education

BS in Computer Science or related field

Tools

JSON
YAML
REST APIs
Jupyter

Job description

Summary

The Apple Intelligence Platform Experience Validation team builds the tooling and automation that keeps Apple Intelligence features high-quality before they ship. We are looking for a Senior SDET to lead the design and implementation of automated model evaluation: standing up LLM-as-a-judge in existing and new pipelines, and building the infrastructure that catches model regressions before they reach human evaluation or the live on population. This is a hands-on, senior individual-contributor role. You will own eval automation as a discipline across the team, partnering with modeling, framework, and infrastructure teams to make model quality a first-class, continuously measured signal.

Description

You will build and maintain model level, component or end-to-end evaluation coverage for the generative features our team validates. Your job is to leverage LLM judge scoring output quality in automation, ensuring reliable, repeatable eval jobs that run that produce actionable signal.

  • Image / visual generation: validating model output and its associated classification metadata, and detecting quality or behavior regressions across model updates.
  • Natural-language generation: evaluating whether generated artifacts and responses match user intent, moving at-desk LLM judges into a scalable and repeatable automation environment.
  • Correctness beyond string matching: replacing exact-match checks for open-ended or factual responses with an LLM-as-judge stage integrated into the pipeline.
  • Generated insights and summaries: assessing whether model-generated content is sensible and good enough to surface to users.
Key Responsibilities
  • Design, build, and maintain LLM-as-a-judge evaluation harnesses and integrate them into existing and new CI/automation pipelines.
  • Author and curate eval sets and rubrics; partner with modeling teams whose own eval sets can run hundreds of examples judged by a separate model.
  • Run eval jobs at scale, triage results, and distinguish real model regressions from rubric problems or infrastructure noise so the signal stays actionable.
  • Plan and build test coverage and tooling from low level component tests through to end-to-end tests that exercise models and frameworks powering Generative AI features.
  • Package eval tooling for reuse — reusable libraries and jobs that other engineers on the team and partner teams can adopt.
  • Define quality gates and reporting so model regressions are caught and communicated before human eval or population rollout.
  • Collaborate with data scientists and modeling engineers on approaches to spot regressions across large output sets or between model updates.
Minimum Qualifications
  • BS in Computer Science, Mathematics, or a related field (or equivalent practical experience)
  • Three years of relevant industry experience in test automation, software development, or related areas.
Preferred Qualifications
  • Strong practical knowledge of Python, including data-pipeline fluency (JSON/YAML, REST APIs).
  • Hands‑on experience with LLM-as-a-judge evaluation and rubric design, or a strong demonstrated ability to ramp into it quickly.
  • Strong software engineering fundamentals — able to define atomic, composable components and build maintainable pipelines and tooling, not just scripts.
  • Strong debugging and triage skills; able to separate genuine regressions from infrastructure or rubric noise.
  • Strong knowledge of the software development lifecycle, testing methodologies, and QA processes.
  • Excellent written and verbal communication; able to document clearly and describe quality signal to modeling and leadership audiences.
  • Ability to lead work across varying priorities and partner multi-functionally with modeling, framework, and infrastructure teams.
  • Experience building on-device tooling and device/model eval infrastructure.
  • Experience integrating with CI/CD and job orchestration systems, and comfort deploying tooling as reusable libraries.
  • Familiarity with generative model behavior — image generation, NLP, or LLM output evaluation.
  • Experience curating and reasoning about large datasets; comfort manually inspecting data (Jupyter or similar) to build intuition and drive next steps.
  • Awareness of dataset bias and fairness considerations in evaluation.
  • Experience with database/query tooling (e.g., SQL) and dashboards/visualization for reporting quality trends.
  • Experience with Xcode is a bonus.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior SDET - LLM Evaluation & Automation for AI Pipelines
Senior SDET - LLM Evaluation & Automation for AI Pipelines

Apple • San Diego (CA)

On-site
USD 140,000 - 190,000
Evaluation & Insights Machine Learning Engineer
Evaluation & Insights Machine Learning Engineer

Apple Inc. • Cupertino (CA)

On-site
USD 150,000 - 230,000
Senior LLM Evaluation Engineer
Senior LLM Evaluation Engineer

Aspire, Jordan • Egypt (PA)

On-site
USD 140,000 - 200,000
AIML - Sr Manager, Evaluation - Data Science & Insights
AIML - Sr Manager, Evaluation - Data Science & Insights

Socket.dev • Seattle (WA)

On-site
USD 180,000 - 240,000
Sr. Evaluation Engineer
Sr. Evaluation Engineer

logicmonitor • San Francisco (CA)

On-site
USD 150,000 - 190,000
Machine Learning Engineer, Software QA - Creativity Apps
Machine Learning Engineer, Software QA - Creativity Apps

Socket.dev • San Diego (CA)

On-site
USD 140,000 - 210,000
Machine Learning - Data Scientist
Machine Learning - Data Scientist

Apple Inc. • Sunnyvale (CA)

On-site
USD 150,400 - 277,600
Comprehensive medical and dental cover
Retirement benefits
Relocation assistance
AIML - Sr Engineering Specialist, Evaluation
AIML - Sr Engineering Specialist, Evaluation

Apple • Seattle (WA)

On-site
USD 130,000 - 180,000
AI Engineer - OH
AI Engineer - OH

LawPro.ai • Kentucky

On-site
USD 150,000 - 190,000
AIML - Sr Applied AI Scientist - GenAI Model Autograding, Evaluation
AIML - Sr Applied AI Scientist - GenAI Model Autograding, Evaluation

Socket.dev • Cupertino (CA)

On-site
USD 180,000 - 240,000