AI Training Infrastructure Engineer - Scale & Performance

SupportFinity™

San Francisco (CA)

On-site

USD 175,000 - 220,000

Full time

14 days+
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Benefits offered by this job

Meaningful equity
Competitive salary
Comprehensive benefits package

Job summary

Fireworks AI is seeking a Member of Technical Staff to design, build, and optimize AI training infrastructure for large-scale models. You’ll collaborate with researchers to create high-performance training pipelines, scale workloads across GPUs and data centers, and implement robust monitoring and storage solutions for massive datasets.

You will contribute to distribution strategies (data/model parallelism, FSDP) and drive automation to improve efficiency, cost, and reliability of training

Qualifications

  • Bachelor's degree or equivalent practical experience.
  • 3+ years of experience with distributed systems and ML infrastructure.
  • Experience with PyTorch.
  • Proficiency in cloud platforms (AWS, GCP, Azure).
  • Experience with containerization, orchestration (Kubernetes, Docker).
  • Knowledge of distributed training techniques (data parallelism, model parallelism, FSDP).

Responsibilities

  • Design and build scalable infrastructure for large-scale model training workloads.
  • Develop and maintain distributed training pipelines for LLMs and multimodal models.
  • Optimize training performance across multiple GPUs, nodes, and data centers.
  • Implement monitoring, logging, and debugging tools for training operations.
  • Architect and maintain data storage solutions for large-scale training datasets.
  • Automate infrastructure provisioning, scaling, and orchestration for model training.
  • Collaborate with researchers to implement and optimize training methodologies.
  • Analyze and improve efficiency, scalability, and cost-effectiveness of training systems.
  • Troubleshoot complex performance issues in distributed training environments.

Skills

Distributed systems
ML infrastructure
PyTorch
Cloud platforms
Kubernetes
Docker
Data parallelism
Model parallelism
FSDP

Education

Bachelor's degree in CS/CE or related field
Master's or PhD (preferred)

Tools

Docker
Kubernetes

Job description

Fireworks AI is seeking a Member of Technical Staff to design, build, and optimize AI training infrastructure for large-scale models. You’ll collaborate with researchers to create high-performance training pipelines, scale workloads across GPUs and data centers, and implement robust monitoring and storage solutions for massive datasets.

You will contribute to distribution strategies (data/model parallelism, FSDP) and drive automation to improve efficiency, cost, and reliability of training

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Distributed AI Training Infrastructure Engineer
Distributed AI Training Infrastructure Engineer

Fireworks AI • United States

Remote
USD 130,000 - 210,000
Performance Optimization Engineer for AI Infrastructure
Performance Optimization Engineer for AI Infrastructure

Fireworks AI • San Mateo (CA)

On-site
USD 180,000 - 260,000
AI Infrastructure Engineer: Scale & Serve Generative AI
AI Infrastructure Engineer: Scale & Serve Generative AI

Fireworks • San Mateo (CA)

On-site
USD 140,000 - 210,000
Research Engineer: AI Training Infrastructure & Innovation
Research Engineer: AI Training Infrastructure & Innovation

Fireworks AI • New York (NY)

On-site
USD 150,000 - 230,000
Senior AI Training Performance Engineer (GPU & Scale)
Senior AI Training Performance Engineer (GPU & Scale)

figure.ai • San Jose (CA), Northern (KY)

Hybrid
USD 200,000 - 400,000
AI Training Performance Engineer — Large-Scale GPU Training
AI Training Performance Engineer — Large-Scale GPU Training

Figure • San Jose (CA)

On-site
USD 200,000 - 400,000
Member of Technical Staff
Member of Technical Staff

Fireworks AI • New York (NY)

On-site
USD 170,000 - 260,000
Member of Technical Staff, Performance Optimization
Member of Technical Staff, Performance Optimization

Fireworks AI • San Mateo (CA)

On-site
USD 180,000 - 260,000
LLM Infrastructure Engineer - Scalable AI Platform
LLM Infrastructure Engineer - Scalable AI Platform

Fireworks AI • San Mateo (CA)

On-site
USD 140,000 - 210,000
Staff AI Runtime Engineer - Scalable GPU Training Platform
Staff AI Runtime Engineer - Scalable GPU Training Platform

Databricks • California (MO)

On-site
USD 190,000 - 265,000