Senior LLM Evaluation Infrastructure Engineer

Inception

San Francisco (CA)

On-site

USD 200,000 - 350,000

Full time

6 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Health insurance
Equity
Flexible vacation
Catered meals
PTO

Job summary

Inception seeks seasoned engineers and scientists to build evaluation infrastructure for frontier LLMs. Youll design frameworks to measure model improvements and ensure reliable production performance.

Join a team of AI researchers and engineers in Palo Alto, contributing to scalable evaluation pipelines, metrics, and benchmarking across tasks, with equity and competitive compensation.

Qualifications

  • BS/MS/PhD in Computer Science, Machine Learning, Statistics, or a related field (or equivalent experience).
  • At least 2 years of experience in ML evaluation, applied ML research, or a related engineering role.
  • Experience with version control (Git), containerization (Docker), and cloud services (AWS, GCP, or Azure).
  • Understanding of LLM fundamentals (autoregressive generation, instruction tuning, RLHF, in-context learning, decoding strategies).
  • Proficiency in Python and ML frameworks such as PyTorch.
  • Experience designing and implementing evaluation metrics and benchmarks for generative models.
  • Excellent communication skills with the ability to distill complex evaluation results into actionable insights.

Responsibilities

  • Build scalable, automated evaluation pipelines that integrate into model training and deployment workflows.
  • Design, develop, and maintain robust evaluation frameworks and benchmarks for measuring LLM performance across diverse tasks and domains.
  • Conduct rigorous statistical analysis of model outputs to identify failure modes, biases, and performance gaps.
  • Partner with product and customer-facing teams to translate real-world use cases into meaningful evaluation criteria.
  • Define and implement quantitative metrics that capture model quality, safety, reliability, and regression detection.

Skills

ML evaluation
Python
PyTorch
Git
Docker
Cloud services
Communication
Statistical analysis

Education

BS/MS/PhD in CS/ML or related field

Tools

Docker
Kubernetes
Terraform

Job description

Inception seeks seasoned engineers and scientists to build evaluation infrastructure for frontier LLMs. Youll design frameworks to measure model improvements and ensure reliable production performance.

Join a team of AI researchers and engineers in Palo Alto, contributing to scalable evaluation pipelines, metrics, and benchmarking across tasks, with equity and competitive compensation.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior LLM Evaluation Scientist
Senior LLM Evaluation Scientist

cohere • New York (NY)

Hybrid
USD 150,000 - 210,000
Lunch stipend
Health benefits
RRSP matching
+5
Senior AI Engineer - Production LLM EvalOps
Senior AI Engineer - Production LLM EvalOps

LawPro.ai • Town of Texas (WI)

On-site
USD 140,000 - 210,000
LLM Training & Evaluation Engineer: Build AI Pipelines
LLM Training & Evaluation Engineer: Build AI Pipelines

Innodata Inc. • United States

On-site
USD 80,000 - 175,000
Senior AI Engineer: LLM Evaluation, Production & Optimization
Senior AI Engineer: LLM Evaluation, Production & Optimization

LawPro.ai • North Carolina

On-site
USD 140,000 - 190,000
Senior AI Engineer: LLM Evaluation & Production
Senior AI Engineer: LLM Evaluation & Production

LawPro.ai • Georgia

On-site
USD 140,000 - 210,000
Senior LLM Evaluation & Post-Training Scientist
Senior LLM Evaluation & Post-Training Scientist

Innodata Inc. • United States

On-site
USD 175,000 - 225,000
Lead AI Engineer — LLM Evaluation & Production Optimizer
Lead AI Engineer — LLM Evaluation & Production Optimizer

LawPro.ai • Kentucky

On-site
USD 150,000 - 190,000
ML Engineer: LLM Evaluation & Observability
ML Engineer: LLM Evaluation & Observability

Gleanwork • Mountain View (CA)

Hybrid
USD 200,000 - 300,000
Health insurance
401(k) plan
Home office improvement stipend
+3
LLM Evaluation Engineering Lead
LLM Evaluation Engineering Lead

DeepRec.ai • Redwood City (CA)

On-site
USD 180,000 - 240,000
High autonomy
Strong technical peers
Meaningful equity
LLM Training & Evaluation Engineer for Scalable AI
LLM Training & Evaluation Engineer for Scalable AI

Innodata Inc. • United States

On-site
USD 80,000 - 175,000