Staff Research Scientist, AI Benchmarking & Evaluation

Thomas Talent Network, LLC

San Francisco (CA)

On-site

USD 140,000 - 210,000

Full time

7 days ago
Be an early applicant
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Job summary

Our client, an AI healthcare software startup, seeks a Member of Technical Staff (Research Scientist) to develop benchmarks and evaluation methodologies for large language models. You will shape how enterprise AI systems are tested before deployment, collaborating with engineering and top AI labs.

Qualifications include 0-3 years in AI/ML research, strong Python, and experience with PyTorch/TF. Familiarity with Django/Flask is a plus; MBA not required.

Qualifications

  • 0-3 years of AI/ML research experience focusing on benchmarking and evaluation.
  • Experience in team environments, including development sprints, Git best practices, PR reviews.
  • Valued: previous AI startup or research lab experience; founder/early employee background.
  • Publications in NLP or benchmarking valued but not required; applied work prioritized.
  • Experience building new benchmarks and evaluation methodologies preferred.
  • Academic background in CS/ML; Masters/PhD preferred; Bachelor's + 0-3 years considered.
  • Background from top tech/AI programs preferred.
  • NLP research experience with publications preferred.
  • Strong Python programming skills for clean, maintainable code.
  • Experience with PyTorch/TF and diffusion models and language modeling.
  • Interest in LLM infrastructure.
  • Experience with Django, Flask, or other Python-based HTTP servers is a plus.
  • Strong communication skills; able to give input and accept feedback.

Responsibilities

  • Evaluate new AI models as they're released.
  • Create fresh benchmarks, hire labelers, construct datasets, and write white papers.
  • Improve auto-evaluation methods for generated text.
  • Collaborate with engineering to implement and scale evaluation methodologies.
  • Work with top AI labs and enterprise customers to understand evaluation needs.

Skills

Python
PyTorch
TensorFlow
NLP research
Benchmarking
Communication
Team collaboration
Research publication

Education

Master's or PhD preferred
Bachelor's degree

Tools

Git
GitHub
Django
Flask

Job description

Our client, an AI healthcare software startup, seeks a Member of Technical Staff (Research Scientist) to develop benchmarks and evaluation methodologies for large language models. You will shape how enterprise AI systems are tested before deployment, collaborating with engineering and top AI labs.

Qualifications include 0-3 years in AI/ML research, strong Python, and experience with PyTorch/TF. Familiarity with Django/Flask is a plus; MBA not required.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Member of Technical Staff
Member of Technical Staff

Thomas Talent Network, LLC • San Francisco (CA)

On-site
USD 140,000 - 210,000
Senior Member of Technical Staff, AI Benchmarking
Senior Member of Technical Staff, AI Benchmarking

Artificial Analysis, Inc. • San Francisco (CA)

On-site
USD 100,000 - 150,000
Competitive compensation including equity
Opportunity to shape AI development
Senior AI Benchmarking & Systems Architect
Senior AI Benchmarking & Systems Architect

Aionia Group • San Francisco (CA)

On-site
USD 130,000 - 220,000
Equity
On-site
Staff AI Benchmark Architect
Staff AI Benchmark Architect

Vals AI, Inc. • San Francisco (CA)

On-site
USD 150,000 - 230,000
Relocation support
Health/dental insurance
Lunch and dinner provided
+2
Remote AI Benchmark & Datasets Engineer
Remote AI Benchmark & Datasets Engineer

Ignite Next GmbH • Palo Alto (CA), Northern (KY)

Hybrid
USD 120,000 - 190,000
Remote work
Office visits Palo Alto, Paris, Wroclw
Head of AI Evaluation & Benchmarks
Head of AI Evaluation & Benchmarks

Vibehackers • San Francisco (CA), Northern (KY)

Hybrid
USD 225,000 - 275,000
Relocation and transportation support
Health and dental insurance
Lunch and dinner provided
+6
AI Benchmarks & Evaluations Program Manager
AI Benchmarks & Evaluations Program Manager

Mercor • San Francisco (CA)

On-site
USD 120,000 - 200,000
Performance bonus structure
Equity grant
$15K relocation bonus
+7
Senior AI Evaluation Scientist — Benchmarks & Systems
Senior AI Evaluation Scientist — Benchmarks & Systems

Oracle • Santa Clara (CA)

On-site
USD 115,000 - 235,000
Medical, dental, and vision insurance
401(k) Savings with company match
Paid time off and holidays
+1
Senior AI Benchmark SME for Quant, Science & Corporate
Senior AI Benchmark SME for Quant, Science & Corporate

Lilt • United States

Remote
USD 90,000 - 130,000
Remote AI Benchmark & Datasets Engineer
Remote AI Benchmark & Datasets Engineer

Pathway Genomics Corporation • Palo Alto (CA)

Remote
USD 150,000 - 210,000
Intellectually stimulating work environment
Work with a pioneering AI startup
Flexible remote work options