AIML - Applied AI Scientist, Image Autograder Systems, Evaluation

Socket.dev

Cupertino (CA)

On-site

USD 150,000 - 210,000

Full time

14 days+
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Job summary

Apple seeks an AI Scientist to join a centralized evaluation org, building autograders for image generation features. You will craft scoring rubrics, train graders and validate outputs at scale, impacting AI experiences for hundreds of millions of users.

You will collaborate with MLEs, data teams and engineers to refine prompts, measure quality, and deploy autograder systems across workflows.

Qualifications

  • Master's or PhD in CS/ML/AI or related field.
  • Deep understanding of visual-language models.
  • Proficiency in Python with production-ready code.
  • Strong communication and stakeholder collaboration skills.

Responsibilities

  • Focus on autograder research, training and adoption.
  • Collaborate with AI feature teams and annotation teams to refine requirements.
  • Develop, evaluate and iterate on grading prompts to align autograder with rubrics.
  • Apply advanced techniques like fine-tuning to close grading gaps.
  • Design analyses to measure and explain autograder quality.
  • Collaborate with MLEs and teams on autograder deployment and adoption.
  • Develop scalable systems to speed up autograder training processes.

Skills

Python
Visual-language models
Communication
Evaluation design

Education

Master's or PhD in Computer Science, Machine Learning, AI or related

Job description

We are looking for an AI Scientist to join a centralized evaluation organization building the next generation of autograders across Apple's most visible image generation AI features. In this role, you will develop autograders that reliably score image output quality at scale. Your work directly influences the quality of AI experiences used by hundreds of millions of Apple customers. This is a high-impact individual contributor role at the intersection of genAI evaluation, genAI model training, data quality, and AI engineering. You will work closely with senior AI scientists, MLEs, data annotation teams, and feature engineers in a fast-moving, technically rigorous environment.

Description
  • In this role you will focus on Autograder research, training and adoption.
  • Collaborate with AI feature teams, eval design teams, and annotation teams to refine feature requirements, grading rubrics and gold annotation sets.
  • Develop, evaluate, and iterate on grading prompts to align autograder behavior with grading rubrics and the gold sets.
  • Identify when prompt tuning reaches its limits and apply other advanced techniques such as fine-tuning to close remaining grading accuracy gaps.
  • Design insightful analysis to measure and explain Autograder quality.
  • Collaborate with MLEs and feature teams on autograder deployment and adoption.
  • Develop scalable system to speed up autograder training processes.
Minimum Qualifications
  • Master's or PhD in Computer Science, Machine Learning, Artificial Intelligence, or a related field.
  • Deep understanding of visual-language models.
  • Familiarity with image quality assessment - perceptual quality dimensions, evaluation metrics, and rubric design.
  • Proficiency in Python; capable of writing well-structured, production-ready model code.
  • Strong communication skills in explaining autograder quality and driving autograder adoption with partner teams.
Preferred Qualifications
  • 1+ years of industry experience in building VLM-based products.
  • Familiarity with autograder and evaluator-specific concepts: grading accuracy, agreement with human raters, calibration, and rubric design.
  • Strong expertise in prompt tuning and fine tuning for VLMs.
  • Demonstrated ability to read AI literature and translate it into applied autograder experiments.
  • Prior experience in building agentic system to scale autograder training/validation.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

AIML - Sr Applied AI Scientist - GenAI Model Autograding, Evaluation
AIML - Sr Applied AI Scientist - GenAI Model Autograding, Evaluation

Socket.dev • Cupertino (CA)

On-site
USD 180,000 - 240,000
Applied AI Scientist – Autograder for Image Quality
Applied AI Scientist – Autograder for Image Quality

Socket.dev • Cupertino (CA)

On-site
USD 150,000 - 210,000
AIML - Applied AI Scientist, Image Autograder Systems, Evaluation
AIML - Applied AI Scientist, Image Autograder Systems, Evaluation

Apple Inc. • Cupertino (CA), Northern (KY)

Hybrid
USD 150,000 - 278,000
Medical and dental coverage
Retirement benefits
Apple stock program/ RSU eligibility
+2
Applied AI Scientist - Image Autograder & Evaluation
Applied AI Scientist - Image Autograder & Evaluation

Apple Inc. • Cupertino (CA)

On-site
USD 150,000 - 278,000
Medical & dental care
Retirement benefits
Discounted products
+3
Applied AI Scientist: Image Autograding at Scale
Applied AI Scientist: Image Autograding at Scale

Apple Inc. • Cupertino (CA), Northern (KY)

Hybrid
USD 150,000 - 278,000
Medical and dental coverage
Retirement benefits
Apple stock program/ RSU eligibility
+2
Autograder AI Scientist: GenAI Quality & Evaluation
Autograder AI Scientist: GenAI Quality & Evaluation

Socket.dev • Cupertino (CA)

On-site
USD 180,000 - 240,000
Autograde AI Scientist for GenAI Quality
Autograde AI Scientist for GenAI Quality

Apple Inc. • Cupertino (CA), Northern (KY)

Hybrid
USD 185,000 - 325,000
Comprehensive medical and dental cover
Retirement benefits
Discounted products and free services
+1
AIML - Sr Applied AI Scientist - GenAI Model Autograding, Evaluation
AIML - Sr Applied AI Scientist - GenAI Model Autograding, Evaluation

Apple Inc. • Cupertino (CA), Northern (KY)

Hybrid
USD 185,000 - 325,000
Comprehensive medical and dental cover
Retirement benefits
Discounted products and free services
+1
Evaluation & Insights Machine Learning Engineer
Evaluation & Insights Machine Learning Engineer

Apple Inc. • Cupertino (CA)

On-site
USD 150,000 - 230,000
AIML - Sr Manager, Evaluation - Data Science & Insights
AIML - Sr Manager, Evaluation - Data Science & Insights

Socket.dev • Seattle (WA)

On-site
USD 180,000 - 240,000