Staff ML Platform Engineer – Large Scale Training (LLMOps/MLOps)

truefoundry

San Francisco (CA)

On-site

USD 180,000 - 280,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Flexible hours
Learning credits
Co-founders mentorship

Job summary

TrueFoundry is seeking a Staff ML Platform Engineer – Large Scale Training to scale deep learning workloads and shipping production-grade solutions.

You will work with PyTorch, multi-node training, and orchestration tools to build a unified compute layer and agentic platform for customers.

Qualifications

  • 5+ years building and deploying ML systems at scale.
  • 5+ years delivering production-grade, high-performance code.
  • Deep experience with multi-GPU/multi-node training (PyTorch preferred).
  • Experience with PyTorch, ML frameworks, and inference engines (vLLM/TensorRT).
  • Kubernetes experience; familiarity with Kubernetes-native tools is a plus.
  • Open-source LLM training/fine-tuning familiarity is a bonus.

Responsibilities

  • Write clean, modular Python code with reliability and performance.
  • Build platform for training and finetuning large-scale ML models across multi-GPU/multi-node clusters.
  • Own infrastructure and code for high-throughput, low-latency inference pipelines.
  • Develop, deploy and evaluate agentic applications for customers.
  • Help shape internal standards for high-scale ML workloads.

Skills

Multi-GPU training
PyTorch
Kubernetes
High-performance code
Distributed systems

Tools

Kubeflow
TensorRT
vLLM

Job description

About TrueFoundry

Every production AI system, whether it's powering customer support, writing code, analyzing financial data, or diagnosing medical conditions, needs the same foundational infrastructure.A way to route between models. A way to manage tools and integrate them securely. A way to orchestrate agents and enforce governance. A unified compute layer to run it all.

That infrastructure layer is being built right now.

We're TrueFoundry, and we're building it. We're looking for a Staff ML Platform Engineer – Large Scale Training (LLMOps/MLOps) to join the team.

The Problem We're Solving

Companies are moving beyond simple chatbots to production agentic systems. These systems route between OpenAI, Anthropic, Google, and self-hosted models. They integrate dozens of tools via protocols like MCP. They orchestrate multi-agent workflows where agents coordinate with other agents.

The infrastructure to support this doesn't exist yet. You can't just duct-tape together a few API calls and call it production-ready.

You need a control plane that handles:

  • Intelligent routing with observability, cost policies, and fallback logic
  • Centralized tool and MCP server management with security and lifecycle controls
  • Agent orchestration with governance and guardrails
  • A unified compute layer to run self-hosted models, custom tools, and agents

We've built two products to solve this:

AI Gateway is the control plane, five composable components (Prompts, LLM Gateway, MCP Gateway, Guardrails, Agent Gateway) that handle routing, orchestration, and governance.

AI Deploy is the compute layer, a Kubernetes-based platform that abstracts ML workloads as standard software primitives, so everything runs on unified infrastructure.

We're Series A, backed by Intel Capital and Sequoia. Companies like CVS, Mastercard, Siemens, Paytm, Synopsys, and Zscaler run production AI workloads on our platform.

We're looking for ML Systems Engineers who are passionate about scaling deep learning workloads, optimizing multi-GPU training, and shipping production-grade solutions. If you live and breathe PyTorch, multi-node training, and love solving gnarly infra challenges—this is your place.

What You’ll Work On
  • Write clean, modular, and scalable Python code, with a strong emphasis on reliability and performance.
  • Build platform for training and finetuning large-scale ML models across multi-GPU, multi-node clusters with PyTorch, Kubeflow, and other orchestration tools.
  • Own the infrastructure and code that enables high-throughput, low-latency inference pipelines for state-of-the-art models.
  • Build platform for developing, deploying and evaluating agentic applications for our end customers.
  • Help shape internal standards and best practices across the engineering team for high-scale ML workloads.
What We’re Looking For
  • 5+ years of hands-on experience building and deploying ML systems at scale.
  • 5+ years of writing production quality high performance code.
  • Deep experience with multi-GPU/multi-node training, ideally with PyTorch as your primary framework.
  • Experience working with torch, high-level ML frameworks, and inference engines (vLLM or TensorRT).
  • Experience with Kubernetes is highly preferred; exposure to Kubernetes-native tools is a huge plus.
  • A pragmatic mindset—you know when to optimize and when to ship.
  • Bonus: Familiarity with open-source LLM training/fine-tuning.
Why Join TrueFoundry?
  • Work directly with ex-Facebook engineers and founders from IIT Kharagpur, UC Berkeley, and Y Combinator alumni.
  • First‑hand exposure to building and scaling a deep-tech startup—insights you’ll carry if you want to start your own one day.
  • Be part of a fearlessly experimental culture focused on customer success and long-term impact.
  • Flexible hours, learning credits, and the opportunity to work shoulder‑to‑shoulder with the co‑founders (Abhishek & Nikunj).
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Staff ML Platform Engineer - Large-Scale Training
Staff ML Platform Engineer - Large-Scale Training

TrueFoundry • San Francisco (CA)

On-site
Staff Engineer, Core Engineering - Flexible Hours
Staff Engineer, Core Engineering - Flexible Hours

TrueFoundry • San Francisco (CA)

On-site
USD 260,000 - 380,000
Staff Engineer – Core Engineering
Staff Engineer – Core Engineering

TrueFoundry • San Francisco (CA)

Remote
USD 260,000 - 380,000
Senior AI/ML Engineer — LLM & Agent Stack (Customer Facing)
Senior AI/ML Engineer — LLM & Agent Stack (Customer Facing)

TrueFoundry • San Mateo (CA)

On-site
USD 150,000 - 210,000
Senior AI/ML Engineer - LLM & Agent Orchestration
Senior AI/ML Engineer - LLM & Agent Orchestration

TrueFoundry • San Mateo (CA)

On-site
USD 150,000 - 210,000
Senior Software Engineer
Senior Software Engineer

TrueFoundry • San Francisco (CA)

Hybrid
USD 180,000 - 240,000
Comprehensive health insurance for you
Lunch and snacks provided
Flex hybrid work: 2 days in office
+1
Staff /Principal Engineer – Core Team
Staff /Principal Engineer – Core Team

TrueFoundry • United States

Hybrid
USD 170,000 - 255,000
Health insurance
401(k) retirement plan
Hybrid work model
+2
Forward Deployed Engineer
Forward Deployed Engineer

TrueFoundry • San Mateo (CA)

Hybrid
USD 150,000 - 190,000
Comprehensive health insurance for you
Flexible hybrid work, 2 days a week in
Lunch and snacks in the office
+1
Forward Deployed Engineer - GTM San Mateo - Engineering
Forward Deployed Engineer - GTM San Mateo - Engineering

Socotra, Inc. • United States

On-site
USD 120,000 - 150,000
Forward Deployed Engineer – GTM
Forward Deployed Engineer – GTM

TrueFoundry • United States

Hybrid
USD 130,000 - 200,000
Health insurance
401(k) retirement plan
Flexible hybrid work (Tue/Wed in the 2
+2