Senior AI Evaluation Scientist — Benchmarks & Systems

Oracle

Santa Clara (CA)

On-site

USD 115,000 - 235,000

Full time

14 days+
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Benefits offered by this job

Medical, dental, and vision insurance
401(k) Savings with company match
Paid time off and holidays
Parental leave

Job summary

Oracle seeks a Senior Applied Scientist to independently own end-to-end evaluations of foundation models, AI agents, and enterprise AI systems. You will translate customer and product needs into testable hypotheses, design benchmarks, and build reproducible evaluation pipelines that inform leadership decisions.

The role requires strong Python skills, experience with ML frameworks, and the ability to communicate complex results clearly to science, engineering, product, and leadership teams.

Qualifications

  • PhD or Master/Bachelor with equivalent experience.
  • Experience designing and executing ML experiments with hypotheses, data, metrics and baselines.
  • Strong knowledge of ML, DL, NLP, and generative AI methods.
  • Hands-on experience evaluating LLMs, foundation models, or ML systems.
  • Proficiency in Python and reliable, maintainable code for experiments.

Responsibilities

  • Independently own end-to-end evaluations of foundation models, AI agents, and enterprise AI systems.
  • Translate customer/product needs into hypotheses, datasets, metrics, baselines and thresholds.
  • Design, implement, and maintain benchmarks and evaluation methods.
  • Write Python production-grade evaluation code and build reproducible pipelines.
  • Analyze model behavior beyond aggregate scores and assess quality, cost, latency, reliability, and safety.

Skills

Python
Machine Learning
Experiment design
Data analysis

Education

PhD in Computer Science, ML, AI, Statistics, Mathematics
Master/Bachelor with equivalent industry experience

Tools

PyTorch
TensorFlow
Hugging Face
NumPy
pandas

Job description

Oracle seeks a Senior Applied Scientist to independently own end-to-end evaluations of foundation models, AI agents, and enterprise AI systems. You will translate customer and product needs into testable hypotheses, design benchmarks, and build reproducible evaluation pipelines that inform leadership decisions.

The role requires strong Python skills, experience with ML frameworks, and the ability to communicate complex results clearly to science, engineering, product, and leadership teams.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior Applied Scientist: AI Evaluation & Benchmarks
Senior Applied Scientist: AI Evaluation & Benchmarks

Oracle • United States

On-site
USD 115,000 - 235,000
Medical, dental, and vision insurance
Short/Long-term disability
Life insurance & AD&D
+2
Senior Foundation-Model Evaluation Scientist
Senior Foundation-Model Evaluation Scientist

Oracle • Seattle (WA)

On-site
USD 115,000 - 235,000
401(k) matching
Employee Stock Purchase Plan
Paid time off
+2
Senior Applied Scientist - End-to-End AI Evaluations
Senior Applied Scientist - End-to-End AI Evaluations

Oracle • Austin (TX)

On-site
USD 115,000 - 234,000
Health insurance
401(k) plan
Paid time off
Senior AI Evaluation Scientist
Senior AI Evaluation Scientist

Oracle • Nashville (TN)

On-site
USD 115,000 - 235,000
Medical, dental, vision insurance
401(k) with company match
Paid time off
Remote Senior AI/ML Benchmarking & Evaluation Engineer
Remote Senior AI/ML Benchmarking & Evaluation Engineer

OpenTeams • Northern (KY)

Hybrid
USD 145,000 - 250,000
Senior Applied Scientist: Research & AI Strategy
Senior Applied Scientist: Research & AI Strategy

Oracle • Seattle (WA)

On-site
USD 158,000 - 355,000
Medical insurance
Dental insurance
Vision insurance
+18
Senior Applied AI Scientist & Research Strategy Lead
Senior Applied AI Scientist & Research Strategy Lead

Oracle • United States

On-site
USD 158,000 - 355,000
Medical, dental, and vision insurance
401(k) Savings and Investment Plan
Paid time off
+2
Staff AI Benchmark Architect
Staff AI Benchmark Architect

Vals AI, Inc. • San Francisco (CA)

On-site
USD 150,000 - 230,000
Relocation support
Health/dental insurance
Lunch and dinner provided
+2
Senior Applied Scientist — AI Strategy & Production
Senior Applied Scientist — AI Strategy & Production

Oracle • Frankfort (KY)

On-site
USD 158,000 - 355,000
Medical Insurance
Dental Insurance
Vision Insurance
+2
AI Evaluations Engineer — Benchmarking Frontiers
AI Evaluations Engineer — Benchmarking Frontiers

Meta • Menlo Park (CA)

On-site
USD 180,000 - 240,000