Senior Applied Scientist

Oracle

United States

On-site

USD 115,000 - 235,000

Full time

10 days ago
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Benefits offered by this job

Medical, dental, and vision insurance
Short/Long-term disability
Life insurance & AD&D
401(k) with match
Paid time off & holidays

Job summary

Oracle's OCI AI Evaluation team seeks a Senior Applied Scientist to own end-to-end model evaluation from question formulation to final recommendations. You will design benchmarks, run experiments on foundation models and AI systems, and produce reproducible evaluation protocols with robust statistical analysis.

You will write production-oriented Python code, handle large datasets, and collaborate with scientific, engineering, product, and leadership teams to deliver actionable results that

Qualifications

  • PhD in Computer Science, ML, AI, Statistics, Mathematics, or equivalent experience; or Master's/Bachelor's with relevant industry background.
  • Experience designing and executing ML experiments, including hypotheses, datasets, metrics, baselines, interpretation of results.
  • Strong knowledge of modern ML, deep learning, NLP, and generative AI methods.
  • Hands-on experience evaluating large language models, foundation models, or ML systems.
  • Proficiency in Python and writing reliable, maintainable code for experiments and data processing.
  • Experience with ML frameworks and data-science libraries (PyTorch, TensorFlow, Hugging Face, NumPy, pandas).
  • Ability to translate technical findings into clear written recommendations for stakeholders.
  • Ability to work independently on complex tasks while collaborating with cross-functional teams.
  • Strong written and verbal communication skills.

Responsibilities

  • Independently own end-to-end evaluations of foundation models and AI systems from question design through analysis and reporting.
  • Translate business needs into evaluation criteria, datasets, metrics, baselines, and acceptance thresholds.
  • Design and maintain benchmarks and evaluation methods across reasoning, coding, RAG, NL2SQL, and more.
  • Write Python evaluation code and build reproducible pipelines and test suites.
  • Evaluate model behavior across quality, cost, latency, safety, robustness, and domain fit.
  • Conduct statistical and error analysis to explain model differences.
  • Develop automated evaluators, including LLM-as-a-judge methods, and validate against human judgments.
  • Design human-evaluation workflows with rubrics, gold datasets, and quality controls.
  • Assess dataset quality, provenance, licensing, privacy, and contamination risks.

Skills

ML concepts
Python
PyTorch
Statistics
Experiment design
Communication

Education

PhD in Computer Science

Tools

PyTorch
TensorFlow
Hugging Face
NumPy
pandas

Job description

Job Description

The OCI AI Evaluation team builds the evidence behind model-selection, product-readiness, and launch decisions. We evaluate frontier foundation models and AI systems across capabilities such as reasoning, coding and agentic coding, retrieval-augmented generation, AI agents, NL2SQL, multimodal understanding, multilingual performance, and responsible AI.

As a Senior Applied Scientist on the team, you will independently own complex evaluation work from problem definition through final recommendation. You will translate ambiguous product and customer questions into measurable hypotheses, select or create appropriate benchmarks, design experiments, build evaluation pipelines, validate data and metrics, analyze failure modes, and communicate conclusions to science, engineering, product, and leadership stakeholders.

This is hands-on applied science. You will write high-quality code, work with large and imperfect datasets, develop and calibrate automated evaluators, and turn one-off analyses into reproducible evaluation protocols and reusable infrastructure. You will examine more than aggregate benchmark scores, considering factors such as statistical validity, data provenance, contamination, robustness, cost, latency, reliability, safety, and operational constraints.

The work sits at the point where research results become product decisions. Success requires scientific rigor, strong engineering judgment, clear writing, and the ability to make progress when requirements, model access, data, or infrastructure are still evolving. You will collaborate closely with other scientists, software engineers, product teams, data and human-annotation teams, and external partners to deliver evaluation results that are technically defensible and useful in practice.

Responsibilities
  • Independently own end-to-end evaluations of foundation models, AI agents, and enterprise AI systems, from initial question and experiment design through analysis, reporting, and stakeholder review.
  • Translate customer, product, and business needs into testable hypotheses, evaluation criteria, datasets, metrics, baselines, and acceptance thresholds.
  • Design, implement, and maintain benchmarks and evaluation methods for areas such as reasoning, coding, agentic workflows, RAG, NL2SQL, multimodal systems, multilingual performance, and responsible AI.
  • Write high-quality Python and production-oriented evaluation code; build reproducible pipelines, test suites, automated checks, and integrations with shared evaluation platforms.
  • Evaluate model and system behavior across quality, cost, latency, reliability, safety, robustness, and domain fit rather than relying only on aggregate scores.
  • Conduct statistical analysis, error analysis, ablations, and qualitative failure-mode investigations to explain model behavior and identify meaningful differences between systems.
  • Develop and validate automated evaluators, including LLM-as-a-judge methods; calibrate them against human judgments and quantify their reliability, bias, and limitations.
  • Design human-evaluation and annotation workflows, including rubrics, gold datasets, sampling plans, quality controls, and vendor or Human-in-the-Loop validation.
  • Assess dataset quality, provenance, representativeness, contamination risk, licensing constraints, privacy, and other factors that could invalidate an evaluation or limit use of its results.
  • Produce concise, decision-ready reports that make methods, assumptions, limitations, tradeoffs, and recommendations explicit for technical and non-technical audiences.
  • Partner with science, engineering, product, data, and operations teams to define requirements, resolve blockers, manage dependencies, and establish clear handoffs and ownership.
  • Turn successful evaluation work into reusable protocols, documented workflows, and shared infrastructure that improve the speed and consistency of future evaluations.
  • Stay current with research in machine learning, generative AI, agent evaluation, and measurement methodology; prototype promising approaches and contribute to science plans, papers, patents, or technical reports where appropriate.
  • Publish original research in top-tier peer-reviewed conferences and journals, and translate relevant evaluation advances into reusable methods, technical reports, or production capabilities for OCI.
  • Review technical work, share expertise, and mentor junior scientists or engineers in experimental design, evaluation methodology, coding, and interpretation of results.
  • Own delivery quality and timelines, communicate risks early, and maintain clear, auditable documentation of experimental configurations, data versions, results, and decisions.
Minimum Qualifications
  • PhD in Computer Science, Machine Learning, Artificial Intelligence, Statistics, Mathematics, or a related quantitative field; or a Master's or Bachelor's degree with equivalent relevant industry experience.
  • Experience designing and executing machine learning experiments, including defining hypotheses, selecting datasets and metrics, establishing baselines, and interpreting results.
  • Strong knowledge of modern machine learning, deep learning, natural language processing, and generative AI methods.
  • Hands-on experience evaluating large language models, foundation models, AI agents, or other machine learning systems.
  • Proficiency in Python and experience writing reliable, maintainable code for experiments, data processing, or machine learning systems.
  • Experience working with machine learning frameworks and data-science libraries such as PyTorch, TensorFlow, Hugging Face, NumPy, pandas, or equivalent tools.
  • Demonstrated ability to analyze complex datasets, identify data-quality problems, and perform statistical, error, and failure-mode analysis.
  • Experience translating technical findings into clear written recommendations for science, engineering, product, or business stakeholders.
  • Ability to work independently on complex assignments while collaborating effectively with scientists, engineers, product managers, and other cross-functional partners.
  • Strong written and verbal communication skills.
Preferred Qualifications
  • Experience building evaluation frameworks, benchmarks, test suites, or observability systems for generative AI models and applications.
  • Experience evaluating one or more of the following: AI agents, agentic coding systems, RAG, NL2SQL, multimodal models, multilingual models, reasoning systems, or responsible AI capabilities.
  • Experience developing or calibrating LLM-as-a-judge methods against human judgments.
  • Experience designing human-evaluation programs, including annotation rubrics, sampling plans, gold datasets, inter-annotator agreement analysis, and quality-control processes.
  • Experience converting research prototypes or one-off experiments into reusable, production-quality tools and workflows.
  • Familiarity with model-evaluation concerns such as benchmark contamination, data provenance, robustness, statistical significance, safety, latency, cost, and reproducibility.
  • Experience deploying or integrating machine learning components in cloud or production environments.
  • Experience working with distributed computing, large-scale datasets, model serving, or parallel evaluation workloads.
  • Publications in top-tier machine learning, natural language processing, data mining, or artificial intelligence conferences or journals.
  • Demonstrated ability to identify new research questions, develop novel evaluation methods, and translate research advances into practical capabilities.
  • Experience mentoring junior scientists or engineers and providing technical or scientific review.
  • Familiarity with enterprise AI requirements, including privacy, security, licensing, compliance, reliability, and auditability.
Qualifications

Disclaimer: Certain U.S. based or U.S. customer or client-facing roles may be required to comply with applicable requirements, such as immunization/occupational health mandates, and/or drug testing requirements.

Range and benefit information provided in this posting are specific to the stated locations only.

US: Hiring Range in USD from: $114,600 - $234,600 per year. May be eligible for bonus, equity, and compensation deferral.

Oracle maintains broad salary ranges for its roles in order to account for variations in knowledge, skills, experience, market conditions and locations, as well as reflect Oracle's differing products, industries and lines of business.

Candidates are typically placed into the range based on the preceding factors as well as internal peer equity.

Oracle US offers a comprehensive benefits package which includes the following:

  • Medical, dental, and vision insurance, including expert medical opinion
  • Short term disability and long term disability
  • Life insurance and AD&D
  • Supplemental life insurance (Employee/Spouse/Child)
  • Health care and dependent care Flexible Spending Accounts
  • Pre-tax commuter and parking benefits
  • 401(k) Savings and Investment Plan with company match
  • Paid time off: Flexible Vacation is provided to all eligible employees assigned to a salaried (non-overtime eligible) position. Accrued Vacation is provided to all other employees eligible for vacation benefits. For employees working at least 35 hours per week, the vacation accrual rate is 13 days annually for the first three years of employment and 18 days annually for subsequent years of employment. Vacation accrual is prorated for employees working between 20 and 34 hours per week. Employees working fewer than 20 hours per week are not eligible for vacation.
  • 11 paid holidays
  • Paid sick leave: 72 hours of paid sick leave upon date of hire. Refreshes each calendar year. Unused balance will carry over each year up to a maximum cap of 112 hours.
  • Paid parental leave
  • Adoption assistance
  • Employee Stock Purchase Plan
  • Financial planning and group legal
  • Voluntary benefits including auto, homeowner and pet insurance

The role will generally accept applications for at least three calendar days from the posting date or as long as the job remains posted.

Career Level - IC3

About Us

Only Oracle brings together the data, infrastructure, applications, and expertise to power everything from industry innovations to life-saving care. And with AI embedded across our products and services, we help customers turn that promise into a better future for all. Discover your potential at a company leading the way in AI and cloud solutions that impact billions of lives.

True innovation starts when everyone is empowered to contribute. That’s why we’re committed to growing a workforce that promotes opportunities for all with competitive benefits that support our people with flexible medical, life insurance, and retirement options. We also encourage employees to give back to their communities through our volunteer programs.

We’re committed to including people with disabilities at all stages of the employment process. If you require accessibility assistance or accommodation for a disability at any point, let us know by emailing accommodation-request_mb@oracle.com or by calling 1-888-404-2494 in the United States.

Oracle is an Equal Employment Opportunity Employer. All qualified applicants will receive consideration for employment without regard to race, color, religion, sex, national origin, sexual orientation, gender identity, disability and protected veterans’ status, or any other characteristic protected by law. Oracle will consider for employment qualified applicants with arrest and conviction records pursuant to applicable law.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior Applied Scientist
Senior Applied Scientist

Oracle • Seattle (WA)

On-site
USD 115,000 - 235,000
401(k) matching
Employee Stock Purchase Plan
Paid time off
+2
Senior Applied Scientist
Senior Applied Scientist

Oracle • Austin (TX)

On-site
USD 115,000 - 234,000
Health insurance
401(k) plan
Paid time off
Senior Applied Scientist
Senior Applied Scientist

Oracle • Santa Clara (CA)

On-site
USD 115,000 - 235,000
Medical, dental, and vision insurance
401(k) Savings with company match
Paid time off and holidays
+1
Senior Applied Scientist
Senior Applied Scientist

Oracle • Nashville (TN)

On-site
USD 115,000 - 235,000
Medical, dental, vision insurance
401(k) with company match
Paid time off
Senior Manager, Platform Software Engineering, AI Acceleration, Developer Tools
Senior Manager, Platform Software Engineering, AI Acceleration, Developer Tools

Oracle • United States

On-site
USD 120,000 - 306,000
Medical, dental, and vision insurance
401(k) with company match
Paid time off
+2
Senior Manager, Platform Software Engineering, AI Acceleration, Developer Tools
Senior Manager, Platform Software Engineering, AI Acceleration, Developer Tools

Oracle • Nashville (TN)

On-site
USD 120,000 - 306,000
Medical, dental, vision insurance
401(k) with company match
Paid time off
Principal Applied Scientist
Principal Applied Scientist

Oracle • United States

On-site
USD 126,000 - 264,000
Medical insurance
Dental insurance
Vision insurance
+8
Senior Core Infrastructure Engineer
Senior Core Infrastructure Engineer

Oracle • Seattle (WA)

On-site
USD 79,000 - 210,000
Medical, dental, and vision insurance
401(k) Savings and Investment Plan
Paid time off and holidays
+1
Principal Applied Scientist
Principal Applied Scientist

Oracle • Seattle (WA)

On-site
USD 126,000 - 264,000
Principal Applied Scientist
Principal Applied Scientist

Oracle • Santa Clara (CA)

On-site
USD 126,000 - 264,000
Medical, dental, and vision insurance
401(k) with company match
Paid time off