Production ML Engineer: Diffusion LLMs

Inception

San Francisco (CA)

On-site

USD 200,000 - 350,000

Full time

3 days ago
Be an early applicant
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Benefits offered by this job

Equity
Health insurance
Dental insurance
Vision insurance
Flexible PTO
Catered meals
Career growth

Job summary

Inception is seeking experienced engineers and researchers to bridge research and real-world application of diffusion LLMs. You will design, train, and deploy models at scale, collaborate with customers to translate their needs into technical solutions, and ensure reliability in production environments.

You will work with cutting-edge tech, from diffusion models to MLOps, and join a fast-growing team in Palo Alto, CA, with competitive compensation, equity, flexible PTO, and comprehensive health

Qualifications

  • BS/MS/PhD in Computer Science, Machine Learning, or a related field (or equivalent experience).
  • At least 5 years of experience working on ML projects in PyTorch (or equivalent).
  • Excellent familiarity with transformers and core LLM concepts (autoregressive pretraining, instruction tuning, in-context learning, LoRA, KV caching).
  • Experience training LLMs, including fine-tuning.
  • Familiarity with large-scale systems and high-performance computing, including GPU/TPU utilization.
  • Experience with version control (Git) and containerization (Docker).
  • Excellent communication skills with the ability to explain technical concepts to non-technical stakeholders.

Responsibilities

  • Design, develop, and optimize our models for production use cases.
  • Partner with customers to understand their requirements and translate them into technical solutions.
  • Implement innovative approaches for post-training generative AI models, including agentic workflows.
  • Work on data preprocessing pipelines, model evaluation, and alignment to enterprise use cases.
  • Contribute to the deployment and maintenance of models in production environments.
  • Collaborate with product teams to design and implement customer-facing ML features.

Skills

PyTorch
Transformers
LLM concepts
Model training
Production systems
Communication skills
GPU/TPU usage

Education

BS/MS/PhD in Computer Science/ML or related field

Tools

Docker
Git
vLLM
TensorRT
Kubernetes

Job description

Inception is seeking experienced engineers and researchers to bridge research and real-world application of diffusion LLMs. You will design, train, and deploy models at scale, collaborate with customers to translate their needs into technical solutions, and ensure reliability in production environments.

You will work with cutting-edge tech, from diffusion models to MLOps, and join a fast-growing team in Palo Alto, CA, with competitive compensation, equity, flexible PTO, and comprehensive health

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Staff ML Systems Engineer - Diffusion LLM Serving
Staff ML Systems Engineer - Diffusion LLM Serving

Inception • San Francisco (CA)

On-site
USD 180,000 - 240,000
Diffusion LLM Research Engineer: Scale & Release
Diffusion LLM Research Engineer: Scale & Release

Institute of Foundation Models • Sunnyvale (CA)

On-site
USD 120,000 - 160,000
Staff Scientist: RL & Reward Modeling for Diffusion LLMs
Staff Scientist: RL & Reward Modeling for Diffusion LLMs

Inception • San Francisco (CA)

On-site
USD 180,000 - 260,000
Staff Scientist — Diffusion LLM Architect
Staff Scientist — Diffusion LLM Architect

Inception • San Francisco (CA)

On-site
USD 180,000 - 230,000
ML Researcher: Diffusion & RL for Creative AI
ML Researcher: Diffusion & RL for Creative AI

Krea • San Francisco (CA), Northern (KY)

Hybrid
USD 180,000 - 260,000
Health & dental insurance
Flexible PTO
401k with company match
+3
Member of Technical Staff, ML Product Engineer
Member of Technical Staff, ML Product Engineer

Inception • San Francisco (CA)

On-site
USD 200,000 - 350,000
Equity
Health insurance
Dental insurance
+4
Remote Inference Engine Engineer - LLMs & Diffusion
Remote Inference Engine Engineer - LLMs & Diffusion

Inferact • United States

Remote
USD 130,000 - 180,000
Inference Runtime Engineer for LLMs & Diffusion
Inference Runtime Engineer for LLMs & Diffusion

Inferact • San Francisco (CA)

Hybrid
USD 200,000 - 400,000
Health, dental, and vision benefits
401(k) company match
Visa sponsorship on case-by-case basis
Senior Backend Engineer, LLM Inference Systems
Senior Backend Engineer, LLM Inference Systems

Inception • San Francisco (CA)

On-site
USD 150,000 - 230,000
Research Engineer - The Diffusion LLM Team
Research Engineer - The Diffusion LLM Team

Institute of Foundation Models • Sunnyvale (CA)

On-site
USD 120,000 - 160,000