Get more replies from employers
Send a job-specific resume in minutes.
Apple seeks a Senior SDET to lead automated model evaluation across our generative AI pipelines, standing up LLM-as-a-judge and integrating it into CI/CD workflows.
You will build end-to-end eval coverage, curate rubrics, run scalable eval jobs, and help surface reliable quality signals before human review or deployment.
This hands-on, senior individual-contributor role requires strong Python, testing, and collaboration with modeling, framework, and infra teams.
The Apple Intelligence Platform Experience Validation team builds the tooling and automation that keeps Apple Intelligence features high-quality before they ship. We are looking for a Senior SDET to lead the design and implementation of automated model evaluation: standing up LLM-as-a-judge in existing and new pipelines, and building the infrastructure that catches model regressions before they reach human evaluation or the live on population. This is a hands-on, senior individual-contributor role. You will own eval automation as a discipline across the team, partnering with modeling, framework, and infrastructure teams to make model quality a first-class, continuously measured signal.
You will build and maintain model level, component or end-to-end evaluation coverage for the generative features our team validates. Your job is to leverage LLM judge scoring output quality in automation, ensuring reliable, repeatable eval jobs that run that produce actionable signal.