Training Infra Engineer for Scalable LLM Systems

Inception

San Francisco (CA)

On-site

USD 180,000 - 240,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Inception is seeking engineers and scientists to design, optimize, and maintain the core systems that enable scalable, efficient training of LLMs. Your work will make experimentation and training fast and reliable so the team can focus on science, not infrastructure bottlenecks.

You'll design distributed training systems across thousands of GPUs, develop high‑performance optimizations, and build reusable frameworks to improve training reproducibility and scalability for new model architectures.

Qualifications

  • BS/MS/PhD in Computer Science, Engineering, or a related field (or equivalent experience).
  • Understanding of ML frameworks (PyTorch, TensorFlow) from a systems perspective.
  • Strong engineering skills — ability to contribute performant, maintainable code and debug in complex codebases.
  • Proficiency in Python and at least one systems programming language (C++/Rust/Go).
  • Experience with containerization (Docker), orchestration (Kubernetes), and CI/CD pipelines.

Responsibilities

  • Design, implement, and optimize distributed training systems that scale across thousands of GPUs and nodes.
  • Develop high-performance optimizations to maximize throughput and efficiency.
  • Develop reusable frameworks and libraries to improve training reproducibility, reliability, and scalability for new model architectures.

Skills

Distributed systems
High performance optimization
Python
C++/Rust/Go
Docker
Kubernetes
ML frameworks: PyTorch, TensorFlow
CI/CD pipelines
Experimentation tooling

Education

BS/MS/PhD in Computer Science/Engineering or related field

Tools

Docker
Kubernetes
Kubeflow
Airflow
Prometheus
Grafana
OpenTelemetry
PyTorch
TensorFlow
PyTorch/XLA
DeepSpeed
Megatron-LM

Job description

Inception is seeking engineers and scientists to design, optimize, and maintain the core systems that enable scalable, efficient training of LLMs. Your work will make experimentation and training fast and reliable so the team can focus on science, not infrastructure bottlenecks.

You'll design distributed training systems across thousands of GPUs, develop high‑performance optimizations, and build reusable frameworks to improve training reproducibility and scalability for new model architectures.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Member of Technical Staff, Training Infra
Member of Technical Staff, Training Infra

Inception • San Francisco (CA)

On-site
USD 180,000 - 240,000
ML Systems Engineer - Scalable Training & Inference
ML Systems Engineer - Scalable Training & Inference

Scale AI, Inc. • New York (NY)

On-site
USD 189,000 - 237,000
Equity
Benefits
Commuter stipend
Tech Lead Manager — LLM Training Platform
Tech Lead Manager — LLM Training Platform

Scale AI, Inc. • New York (NY)

On-site
USD 264,000 - 331,000
Health, dental and vision coverage
Retirement benefits
Learning and development stipend
+2
LLM Pre-training & Distributed Engineer (AI Infrastructure)
LLM Pre-training & Distributed Engineer (AI Infrastructure)

Hyphen Connect Limited • San Francisco (CA)

On-site
USD 120,000 - 160,000
LLM Pre-training & Distributed Engineer (AI Infrastructure)
LLM Pre-training & Distributed Engineer (AI Infrastructure)

Hyphen Connect Limited • Oregon (WI)

On-site
USD 100,000 - 130,000
LLM Training & Inference Systems Engineer
LLM Training & Inference Systems Engineer

United States Digital Space LLC • New York (NY), San Francisco (CA)

On-site
USD 190,000 - 237,000
Staff ML Systems Engineer - Diffusion LLM Serving
Staff ML Systems Engineer - Diffusion LLM Serving

Inception • San Francisco (CA)

On-site
USD 180,000 - 240,000
Staff ML Infra Engineer: Distributed Training & Inference
Staff ML Infra Engineer: Distributed Training & Inference

Jobtailor • Boston (MA)

On-site
USD 120,000 - 160,000
ML Infra Engineer — GPU Clusters & Distributed Systems
ML Infra Engineer — GPU Clusters & Distributed Systems

AI Breaking Wire • San Francisco (CA), Northern (KY)

Hybrid
USD 170,000 - 250,000
Industry-leading compensation and/or:?
Unlimited PTO
Top-tier medical, dental, and vision
+1
Staff Engineer, Scalable RL Infrastructure
Staff Engineer, Scalable RL Infrastructure

Inception • San Francisco (CA)

On-site
USD 180,000 - 240,000