Staff Engineer, Distributed ML Infrastructure

Fireworks AI

United States

On-site

USD 140,000 - 220,000

Full time

9 days ago
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Job summary

Fireworks AI is hiring a Training Infrastructure Engineer to design, build, and maintain large-scale backend and cloud-native infrastructure for distributed ML training, inference, and data processing.

You will lead technical design discussions, mentor engineers, and drive best practices for scalable ML systems, focusing on reliability, low latency, and cost optimization. Proficiency with Docker, Kubernetes, and modern server-side languages is required.

Qualifications

  • Design, develop, and optimize large-scale backend infrastructure and distributed data systems in cloud environments.
  • Define and implement data-driven metrics to support company goals.
  • Lead cross-functional projects and contribute to technical design documentation.
  • Conduct coding interviews and provide structured feedback for engineering candidates.
  • Collaborate with ML, DevOps, and product teams to translate requirements into infrastructure solutions.
  • Experience with cloud-native tooling and infrastructure (Docker, Kubernetes).
  • Proficiency in major server-side languages (Python, C++, Go, TypeScript) and data processing/API systems (gRPC/Thrift).

Responsibilities

  • Design, build, and maintain scalable backend infrastructure for ML training, inference, and data processing pipelines.
  • Architect scalable, resilient systems emphasizing reliability and low latency.
  • Lead design discussions, mentor engineers, and establish best practices for ML infrastructure.
  • Optimize compute, storage, and network performance to reduce costs and improve efficiency.
  • Collaborate with ML, DevOps, and product teams to convert research requirements into robust infra.
  • Evaluate and integrate cloud-native/open-source tech (Kubernetes, Ray, Kubeflow, MLFlow).
  • Own end-to-end systems from design to deployment with fault tolerance and operations in mind.

Skills

Backend design
Distributed data systems
AB testing
Cross-functional collaboration
Mentoring engineers
Interviewing engineers

Education

Bachelor's degree in Computer Science or related field

Tools

Docker
Kubernetes
Python
Go
TypeScript
C++
PostgreSQL
MySQL
DynamoDB
Apache Spark
Apache Kafka
Apache Flink
Ray
Kubeflow
MLFlow
gRPC
Thrift

Job description

Fireworks AI is hiring a Training Infrastructure Engineer to design, build, and maintain large-scale backend and cloud-native infrastructure for distributed ML training, inference, and data processing.

You will lead technical design discussions, mentor engineers, and drive best practices for scalable ML systems, focusing on reliability, low latency, and cost optimization. Proficiency with Docker, Kubernetes, and modern server-side languages is required.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Cloud-Native Training Infrastructure Engineer
Cloud-Native Training Infrastructure Engineer

Fireworks AI • New York (NY)

On-site
USD 170,000 - 260,000
Distributed AI Training Infrastructure Engineer
Distributed AI Training Infrastructure Engineer

Fireworks AI • United States

Remote
USD 130,000 - 210,000
Senior Cloud Infrastructure Engineer for ML Platforms
Senior Cloud Infrastructure Engineer for ML Platforms

Fireworks AI • United States

On-site
USD 180,000 - 260,000
AI Training Infrastructure Engineer — Distributed Systems
AI Training Infrastructure Engineer — Distributed Systems

Fireworks AI • New York (NY)

On-site
USD 140,000 - 220,000
Senior Cloud Infrastructure Engineer - Scalable ML Platform
Senior Cloud Infrastructure Engineer - Scalable ML Platform

Fireworks • New York (NY), San Mateo (CA)

On-site
USD 140,000 - 210,000
Member of Technical Staff, AI Training Infrastructure
Member of Technical Staff, AI Training Infrastructure

Fireworks AI • New York (NY)

On-site
USD 140,000 - 220,000
LLM Infrastructure Engineer - Scalable AI Platform
LLM Infrastructure Engineer - Scalable AI Platform

Fireworks AI • San Mateo (CA)

On-site
USD 140,000 - 210,000
Staff Backend Architect, Multi-Cloud AI Infra & Schedulers
Staff Backend Architect, Multi-Cloud AI Infra & Schedulers

Fireworks AI • San Mateo (CA)

On-site
USD 180,000 - 240,000
AI Training Infrastructure Engineer - Scale LLM Training
AI Training Infrastructure Engineer - Scale LLM Training

Fireworks AI • San Mateo (CA)

On-site
USD 175,000 - 220,000
Equity
Comprehensive benefits
Competitive salary
Staff ML Infrastructure Architect
Staff ML Infrastructure Architect

Adobe Inc. • San Jose (CA)

On-site
USD 180,000 - 260,000