Senior Applied Scientist: AI Evaluation & Benchmarks

Oracle

United States

On-site

USD 115,000 - 235,000

Full time

28 hours ago
Be an early applicant
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Benefits offered by this job

Medical, dental, and vision insurance
Short/Long-term disability
Life insurance & AD&D
401(k) with match
Paid time off & holidays

Job summary

Oracle's OCI AI Evaluation team seeks a Senior Applied Scientist to own end-to-end model evaluation from question formulation to final recommendations. You will design benchmarks, run experiments on foundation models and AI systems, and produce reproducible evaluation protocols with robust statistical analysis.

You will write production-oriented Python code, handle large datasets, and collaborate with scientific, engineering, product, and leadership teams to deliver actionable results that

Qualifications

  • PhD in Computer Science, ML, AI, Statistics, Mathematics, or equivalent experience; or Master's/Bachelor's with relevant industry background.
  • Experience designing and executing ML experiments, including hypotheses, datasets, metrics, baselines, interpretation of results.
  • Strong knowledge of modern ML, deep learning, NLP, and generative AI methods.
  • Hands-on experience evaluating large language models, foundation models, or ML systems.
  • Proficiency in Python and writing reliable, maintainable code for experiments and data processing.
  • Experience with ML frameworks and data-science libraries (PyTorch, TensorFlow, Hugging Face, NumPy, pandas).
  • Ability to translate technical findings into clear written recommendations for stakeholders.
  • Ability to work independently on complex tasks while collaborating with cross-functional teams.
  • Strong written and verbal communication skills.

Responsibilities

  • Independently own end-to-end evaluations of foundation models and AI systems from question design through analysis and reporting.
  • Translate business needs into evaluation criteria, datasets, metrics, baselines, and acceptance thresholds.
  • Design and maintain benchmarks and evaluation methods across reasoning, coding, RAG, NL2SQL, and more.
  • Write Python evaluation code and build reproducible pipelines and test suites.
  • Evaluate model behavior across quality, cost, latency, safety, robustness, and domain fit.
  • Conduct statistical and error analysis to explain model differences.
  • Develop automated evaluators, including LLM-as-a-judge methods, and validate against human judgments.
  • Design human-evaluation workflows with rubrics, gold datasets, and quality controls.
  • Assess dataset quality, provenance, licensing, privacy, and contamination risks.

Skills

ML concepts
Python
PyTorch
Statistics
Experiment design
Communication

Education

PhD in Computer Science

Tools

PyTorch
TensorFlow
Hugging Face
NumPy
pandas

Job description

Oracle's OCI AI Evaluation team seeks a Senior Applied Scientist to own end-to-end model evaluation from question formulation to final recommendations. You will design benchmarks, run experiments on foundation models and AI systems, and produce reproducible evaluation protocols with robust statistical analysis.

You will write production-oriented Python code, handle large datasets, and collaborate with scientific, engineering, product, and leadership teams to deliver actionable results that

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Applied Scientist - End-to-End AI Evaluations
Senior Applied Scientist - End-to-End AI Evaluations

Oracle • Austin (TX)

On-site
USD 115,000 - 234,000
Health insurance
401(k) plan
Paid time off
Senior Foundation-Model Evaluation Scientist
Senior Foundation-Model Evaluation Scientist

Oracle • Seattle (WA)

On-site
USD 115,000 - 235,000
401(k) matching
Employee Stock Purchase Plan
Paid time off
+2
Senior Applied Scientist
Senior Applied Scientist

Ll Oefentherapie • Denver (CO)

On-site
USD 150,000 - 210,000
Senior Applied Scientist: Research & AI Strategy
Senior Applied Scientist: Research & AI Strategy

Oracle • Seattle (WA)

On-site
USD 158,000 - 355,000
Medical insurance
Dental insurance
Vision insurance
+18
Senior Principal AI Platform Scientist
Senior Principal AI Platform Scientist

Oracle • Nashville (TN)

On-site
USD 158,000 - 355,000
Health benefits
401(k) match
Paid time off
Senior Applied Scientist – AI Research & Production
Senior Applied Scientist – AI Research & Production

Oracle • Austin (TX)

On-site
USD 126,000 - 264,000
Medical, dental, and vision insurance
401(k) Savings and Investment Plan
Paid time off
+2
Senior Applied AI Scientist & Research Strategy Lead
Senior Applied AI Scientist & Research Strategy Lead

Oracle • United States

On-site
USD 158,000 - 355,000
Medical, dental, and vision insurance
401(k) Savings and Investment Plan
Paid time off
+2
Senior Principal AI Scientist - Enterprise Cloud Platform
Senior Principal AI Scientist - Enterprise Cloud Platform

Oracle • United States

On-site
USD 158,000 - 355,000
Medical insurance
401(k) with company match
Paid time off
Senior AI/ML Evaluation Engineer — Benchmarks (Remote)
Senior AI/ML Evaluation Engineer — Benchmarks (Remote)

OpenTeams • Washington, Denver (CO), Colorado Springs (CO)

Hybrid
USD 145,000 - 250,000
401(k) Match – Up to 5% with full vest
Unlimited PTO – 15 days minimum
Fully Remote Setup – up to $3,000 for
+3
Senior Applied Scientist
Senior Applied Scientist

Oracle • Seattle (WA)

On-site
USD 115,000 - 235,000
401(k) matching
Employee Stock Purchase Plan
Paid time off
+2