Member of Technical Staff, Training Infra

Inception

San Francisco (CA)

On-site

USD 180,000 - 240,000

Full time

14 days+
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Job summary

Inception is seeking engineers and scientists to design, optimize, and maintain the core systems that enable scalable, efficient training of LLMs. Your work will make experimentation and training fast and reliable so the team can focus on science, not infrastructure bottlenecks.

You'll design distributed training systems across thousands of GPUs, develop high‑performance optimizations, and build reusable frameworks to improve training reproducibility and scalability for new model architectures.

Qualifications

  • BS/MS/PhD in Computer Science, Engineering, or a related field (or equivalent experience).
  • Understanding of ML frameworks (PyTorch, TensorFlow) from a systems perspective.
  • Strong engineering skills — ability to contribute performant, maintainable code and debug in complex codebases.
  • Proficiency in Python and at least one systems programming language (C++/Rust/Go).
  • Experience with containerization (Docker), orchestration (Kubernetes), and CI/CD pipelines.

Responsibilities

  • Design, implement, and optimize distributed training systems that scale across thousands of GPUs and nodes.
  • Develop high-performance optimizations to maximize throughput and efficiency.
  • Develop reusable frameworks and libraries to improve training reproducibility, reliability, and scalability for new model architectures.

Skills

Distributed systems
High performance optimization
Python
C++/Rust/Go
Docker
Kubernetes
ML frameworks: PyTorch, TensorFlow
CI/CD pipelines
Experimentation tooling

Education

BS/MS/PhD in Computer Science/Engineering or related field

Tools

Docker
Kubernetes
Kubeflow
Airflow
Prometheus
Grafana
OpenTelemetry
PyTorch
TensorFlow
PyTorch/XLA
DeepSpeed
Megatron-LM

Job description

The Role

We\'re looking for engineers and scientists to design, optimize, and maintain the core systems that enable scalable, efficient training of LLMs. Your goal is to make experimentation and training at Inception fast and reliable so our team can focus on science, not system bottlenecks.


Key Responsibilities


  • Design, implement, and optimize distributed training systems that scale across thousands of GPUs and nodes.

  • Develop high-performance optimizations to maximize throughput and efficiency.

  • Develop reusable frameworks and libraries to improve training reproducibility, reliability, and scalability for new model architectures.


Qualifications


  • BS/MS/PhD in Computer Science, Engineering, or a related field (or equivalent experience).

  • Understanding of ML frameworks (PyTorch, TensorFlow) from a systems perspective.

  • Strong engineering skills — ability to contribute performant, maintainable code and debug in complex codebases.

  • Proficiency in Python and at least one systems programming language (C++/Rust/Go).

  • Experience with containerization (Docker), orchestration (Kubernetes), and CI/CD pipelines.


Preferred Skills


  • Experience building and maintaining large-scale language models with tens of billions of parameters or more.

  • Experience with ML workflow orchestration tools (Kubeflow, Airflow).

  • Background in performance optimization and profiling of ML systems (Prometheus, Grafana, OpenTelemetry).

  • Familiarity with distributed frameworks such as PyTorch/XLA, DeepSpeed, Megatron-LM.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Training Infra Engineer for Scalable LLM Systems
Training Infra Engineer for Scalable LLM Systems

Inception • San Francisco (CA)

On-site
USD 180,000 - 240,000
Member of Technical Staff — Training Infrastructure
Member of Technical Staff — Training Infrastructure

Kindredventures • San Francisco (CA)

On-site
USD 180,000 - 240,000
Member of Technical Staff, Inference & Serving
Member of Technical Staff, Inference & Serving

Inception • San Francisco (CA)

On-site
USD 180,000 - 240,000
Member of Technical Staff, Data Infrastructure
Member of Technical Staff, Data Infrastructure

Inception • San Francisco (CA)

On-site
USD 140,000 - 190,000
Member of Technical Staff - ML Infra
Member of Technical Staff - ML Infra

Kindredventures • San Francisco (CA)

On-site
USD 160,000 - 220,000
Member of Technical Staff, MLSys
Member of Technical Staff, MLSys

Bake AI • San Mateo (CA), Northern (KY)

Hybrid
USD 180,000 - 240,000
LLM Pre-training & Distributed Engineer (AI Infrastructure)
LLM Pre-training & Distributed Engineer (AI Infrastructure)

Hyphen Connect Limited • Oregon (WI)

On-site
USD 100,000 - 130,000
LLM Pre-training & Distributed Engineer (AI Infrastructure)
LLM Pre-training & Distributed Engineer (AI Infrastructure)

Hyphen Connect Limited • San Francisco (CA)

On-site
USD 120,000 - 160,000
Tech Lead Manager for Scalable LLM Training Platform
Tech Lead Manager for Scalable LLM Training Platform

United States Digital Space LLC • San Francisco (CA), New York (NY)

On-site
USD 290,000 - 363,000
Software Engineer - ML Infrastructure
Software Engineer - ML Infrastructure

Epsilon • San Francisco (CA), Northern (KY)

Hybrid
USD 180,000 - 280,000