Senior Research Scientist, STEM

Turing

Palo Alto (CA)

On-site

USD 250,000 - 350,000

Full time

2 days ago
Be an early applicant
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Benefits offered by this job

Equity

Job summary

Turing seeks a Senior Research Scientist to advance frontier benchmarks, synthetic data, and evaluation methodologies. You will propose bold research directions, design rigorous experiments, and publish results that shape the field.

The role emphasizes evaluating frontier STEM capabilities, reducing hallucination, and enabling reliable, agentic scientific work. Office-based, five days a week, in San Francisco, Palo Alto, or Seattle, with compensation ranging from $250,000 to $350,000 plus equity.

Qualifications

  • PhD or equivalent research experience in a highly technical field.
  • Demonstrated ability to formulate and execute original research.
  • Strong understanding of modern LLMs and frontier AI research.
  • Excellent experimental design, quantitative reasoning and judgment.
  • Strong Python skills to build research prototypes and evaluation pipelines.

Responsibilities

  • Identify high-impact gaps in benchmark and evaluation literature.
  • Design novel benchmarks in STEM fields and model functionality.
  • Develop evaluations for emerging model capabilities.
  • Design rigorous task-generation, grading, contamination-control and validation.
  • Build benchmarks that become valuable research contributions.
  • Develop methods for generating high-quality synthetic STEM training data.
  • Study how data quality affects downstream performance.
  • Study hallucination, uncertainty, calibration, and verification.
  • Investigate agentic science workflows and collaboration.
  • Propose new research directions and programs at Turing.

Skills

Python
Experimental design
Scientific writing
Data analysis
Communication

Education

PhD or equivalent research experience

Job description

About Turing

Turing’s mission is to accelerate superintelligence to drive real economic progress. Headquartered in San Francisco, Turing works with frontier AI labs to generate high-quality datasets, reinforcement learning environments, and frontier research benchmarks that improve model capabilities in software engineering, enterprise knowledge work, and advanced STEM reasoning. In software engineering, Turing is the largest and longest-running data provider in the category. Turing also works with Fortune 500 enterprises across financial services, life sciences, healthcare, retail, automotive, and CPG to build and deploy end-to-end agentic AI systems inside mission-critical workflows. By operating on both sides, Turing closes the loop between frontier research and enterprise deployment, turning real-world deployment signals into better data, evaluations, and more capable models. Learn more at www.turing.com.

The Role

Turing is seeking exceptional Senior Research Scientists to join our STEM research organization and develop new ways to evaluate, train, and improve frontier AI systems.

This is a research-first role focused on problems where the right benchmark, dataset, or methodology often does not yet exist. You will identify important gaps in the literature, propose ambitious new research directions, and take projects from initial hypothesis through experimentation, benchmark construction, and publication.

Our research is deliberately focused on frontier STEM evaluation, synthetic data, hallucination and reliability, and agentic science. We are looking for scientists who can recognize important problems early, formulate them precisely, and design rigorous research programs to answer them.

What You’ll Do
Frontier benchmarks and evaluation
  • Identify high-impact gaps in existing benchmark and evaluation literature.
  • Design novel benchmarks in and across STEM fields and on general model functionality.
  • Develop evaluations for emerging model capabilities that are poorly captured by traditional static benchmarks.
  • Design rigorous task-generation, grading, contamination-control, difficulty-calibration, and validation methodologies.
  • Build benchmarks that can become both valuable research contributions and meaningful standards for evaluating frontier models.
Synthetic data and post-training
  • Develop methods for generating high-quality synthetic STEM training data.
  • Study how task selection, difficulty, diversity, verification, filtering, and data quality affect downstream performance.
  • Explore methods for generating useful training signal in domains where expert human data is scarce or expensive.
  • Design experiments that determine when synthetic data genuinely improves capabilities rather than simply increasing training volume.
Hallucination, reliability, and verification
  • Study hallucination, uncertainty, calibration, and epistemic failure in technical domains.
  • Develop evaluations and methods for improving factual reliability, self-correction, verification, citation, and appropriate abstention.
  • Investigate when models should reason internally, invoke tools, seek external evidence, or recognize that they do not know.
Agentic science
  • Research AI systems capable of performing extended scientific and technical work.
  • Develop workflows involving literature search, coding, simulation, tool use, experimentation, verification, and iterative reasoning.
  • Evaluate long-horizon scientific agents and identify the bottlenecks preventing them from reliably performing real research.
  • Explore new approaches to human-AI and multi-agent scientific collaboration.
New research directions

The areas above are our core focus, not an exhaustive list. Researchers will also have significant latitude to propose new programs in areas such as reasoning, model evaluation, AI-for-science, data generation, and emerging capabilities.

What We’re Looking For
  • PhD or equivalent research experience in machine learning, computer science, mathematics, physics, chemistry, biology, engineering, statistics, or another highly technical field.
  • Demonstrated ability to formulate and execute original research.
  • Strong understanding of modern LLMs and the frontier AI research landscape.
  • Excellent experimental design, quantitative reasoning, and scientific judgment.
  • Ability to rapidly understand unfamiliar technical literature and develop expertise in new areas.
  • Strong Python skills and the ability to independently build research prototypes and evaluation pipelines.
  • Excellent technical writing and communication.
  • Comfort working in a fast-moving environment where the research agenda evolves with the frontier.

A strong publication record is valuable, but we care most about whether you can identify important questions, design rigorous ways to answer them, and execute quickly enough for the results to matter.

What Success Looks Like

You might:

  • Identify a major capability that existing benchmarks fail to measure and create the benchmark that becomes the standard for evaluating it.
  • Discover a failure mode in current synthetic-data pipelines and develop a method that materially improves post-training.
  • Build a new evaluation that changes how frontier labs understand hallucination, reasoning, or scientific capability.
  • Develop an agentic workflow that substantially advances performance on complex scientific research tasks.
  • Launch an entirely new research direction that grows into a major program within Turing.
Why Turing

Frontier models are improving faster than the benchmarks, datasets, and research methodologies used to understand them.

The STEM research team at Turing works on that gap directly. Our goal is not simply to apply existing evaluation methods, but invent the benchmarks, data-generation methods, and research frameworks needed for the next generation of AI systems.

If you want to define how frontier AI is evaluated and improved across science and technical reasoning, we’d like to hear from you.

This role is required to be in office five days a week, based in any of Turing’s offices in San Francisco, Palo Alto, or Seattle.

Compensation: $250,000 to $350,000 OTE + Equity

Values
  • We are client first: We put our clients at the center of everything we do, because their success is the ultimate measure of our value.

  • We work at Start-Up Speed: We move fast, stay agile and favor action because momentum is the foundation of perfection

  • We are AI forward: We help our clients build the future of Al and implement it in our own roles and workflow to amplify productivity.

Advantages of joining Turing
  • Work at the frontier of AI , helping the world’s leading AI labs improve their most advanced models by building expert datasets, RL environments, and first-of-a-kind benchmarks.

  • Contribute to leading-edge AI research and showcase your work at top conferences such as ICLR, ICML, and NeurIPS.

  • Bring frontier AI innovation to the enterprise , applying lessons learned from leading AI labs to solve real-world business challenges.

  • Collaborate with and learn from exceptional colleagues with deep AI experience from Google, Meta, Amazon, and other leading technology companies.

  • Move at the pace of AI innovation , with the speed, ownership, and impact of a startup.

Turing is proud to be an equal opportunity employer. We do not discriminate on the basis of race, religion, color, national origin, gender, gender identity, sexual orientation, age, marital status, disability, protected veteran status, or any other legally protected characteristics. At Turing we are dedicated to building a diverse, inclusive and authentic workplaceand celebrate authenticity, so if you’re excited about this role but your past experience doesn’t align perfectly with every qualification in the job description, we encourage you to apply anyways. You may be just the right candidate for this or other roles.

For applicants from the European Union, please review Turing’s GDPR notice here.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Staff Research Scientist, STEM
Staff Research Scientist, STEM

Turing • Palo Alto (CA)

On-site
USD 250,000 - 400,000
Equity
Research Scientist, STEM
Research Scientist, STEM

Turing • Palo Alto (CA), San Francisco (CA), Seattle (WA)

On-site
USD 150,000 - 300,000
Equity
Senior Research Scientist, STEM
Senior Research Scientist, STEM

turing • San Francisco (CA)

On-site
USD 250,000 - 350,000
Senior Research Engineer, Enterprise Knowledge
Senior Research Engineer, Enterprise Knowledge

Turing • Palo Alto (CA)

On-site
USD 250,000 - 350,000
Senior Research Engineer, Frontier Data
Senior Research Engineer, Frontier Data

Turing • Palo Alto (CA), San Francisco (CA), Seattle (WA)

On-site
USD 250,000 - 350,000
Equity
Staff Research Engineer, Frontier Data
Staff Research Engineer, Frontier Data

Turing • Palo Alto (CA), San Francisco (CA), Seattle (WA)

On-site
USD 250,000 - 400,000
Work at frontier of AI
Conferences publications opportunities
Equity
Senior Research Engineer
Senior Research Engineer

turing • Palo Alto (CA)

On-site
USD 250,000 - 350,000
Equity
Office in SF/Palo Alto/Seattle
Senior Research Engineer
Senior Research Engineer

Cerebras • California (MO)

On-site
USD 160,000 - 230,000
Work culture
Colleagues
Compensation
+1
Strategic Project Lead, Software Engineering
Strategic Project Lead, Software Engineering

Cerebras • United States

On-site
USD 120,000 - 200,000
Flexible working hours
Competitive compensation
Amazing work culture
Strategic Project Lead, Software Engineering
Strategic Project Lead, Software Engineering

Turing • New York (NY)

On-site
USD 120,000 - 280,000
Amazing work culture
Awesome colleagues
Competitive compensation