Lead ML Model Training & Post-Training Architect

Inflection AI, Inc.

Palo Alto, Northern (CA, KY)

Hybrid

USD 400,000 - 550,000

Full time

9 days ago

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Robust medical, dental and vision with
401k matching
Flexible time off
Holidays & leave
Equity participation
Visa support for Bay Area

Job summary

Inflection AI, Inc. in the San Francisco Bay Area is seeking a Principal Research Engineer to own the model-improvement loop from data and training through evals, post-training, release criteria and production feedback.

This hands-on technical leader will drive training strategies, architecture decisions, and large-scale distributed training across GPUs to ship models that are measurably better for users.

Qualifications

  • Experience leading large-scale LLM or foundation-model training or post-training programs.
  • Strong experience with transformer-based models and distributed training systems.
  • Experience operating large-scale training infrastructure, e.g., GPU clusters.

Responsibilities

  • Own the model-improvement roadmap across capability, reliability, and enterprise readiness.
  • Lead training and post-training strategy, including supervised fine-tuning and reward modeling.
  • Drive architecture and optimization decisions for training and inference.
  • Coordinate large-scale training on distributed GPU clusters (thousands of GPUs).
  • Define data strategy and evaluation data for production use and feedback loops.
  • Build and improve evaluation and release-quality systems and model-readiness reviews.
  • Partner with infrastructure and research teams to improve reliability and cost-performance.
  • Debug and improve model behavior across data, training, and production.

Skills

Transformer models
Distributed training
Post-training methods
RLHF
SFT
DPO
GRPO
RLAIF
Model evaluation
Budgeting cost and throughput

Education

PhD in Computer Science / ML / AI or related field

Tools

PyTorch
TensorFlow

Job description

Inflection AI, Inc. in the San Francisco Bay Area is seeking a Principal Research Engineer to own the model-improvement loop from data and training through evals, post-training, release criteria and production feedback.

This hands-on technical leader will drive training strategies, architecture decisions, and large-scale distributed training across GPUs to ship models that are measurably better for users.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior AI/ML Engineer
Senior AI/ML Engineer

Clera • San Francisco (CA)

On-site
USD 130,000 - 160,000
Principal AI Model Training & Post-Training Lead
Principal AI Model Training & Post-Training Lead

Inflection AI • Palo Alto (CA)

On-site
USD 400,000 - 550,000
Diverse medical, dental and vision options
401k matching program
Unlimited paid time off
+2
Lead ML Infrastructure & Evaluation
Lead ML Infrastructure & Evaluation

Cursor • New York (NY)

On-site
USD 180,000 - 260,000
Senior ML Architect - Build Scalable Production Models
Senior ML Architect - Build Scalable Production Models

European Recruitment BV • United States

On-site
USD 180,000 - 230,000
Lead ML Systems Engineer: Production-Grade AI
Lead ML Systems Engineer: Production-Grade AI

A1 • Palo Alto (CA)

On-site
USD 190,000 - 230,000
Senior ML Engineer: Edge AI & Post-Training Pipelines
Senior ML Engineer: Edge AI & Post-Training Pipelines

Intel • Folsom (CA)

Hybrid
USD 195,000 - 361,000
Hybrid work model
Stock bonuses
Health benefits
Production ML Engineer — Scale & Train LLMs
Production ML Engineer — Scale & Train LLMs

Anthropic • San Francisco (CA)

On-site
USD 350,000 - 850,000
Generous vacation and parental leave
Flexible working hours
Office space for collaboration
Lead ML Platform Engineer: Training & Inference at Scale
Lead ML Platform Engineer: Training & Inference at Scale

Paramount • Burbank (CA)

On-site
USD 157,000 - 235,000
Benefits package
On-site & virtual events
Generous PTO
Staff ML Engineer: End-to-End Model Training & Deployment
Staff ML Engineer: End-to-End Model Training & Deployment

Parallel Bio • San Francisco (CA)

On-site
USD 100,000 - 140,000
Competitive salary
Generous equity
Visa sponsorships
+5
Lead ML Engineer, Physics-Informed AI & Digital Twins
Lead ML Engineer, Physics-Informed AI & Digital Twins

Stand • San Francisco (CA)

On-site
USD 250,000 - 295,000
Above-market Health, Dental, and Vision coverage
Weekly lunch stipend
Flexible time off + holidays
+6