Distributed AI Training Infrastructure Engineer

Fireworks AI

United States

Remote

USD 130,000 - 210,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Fireworks AI in the United States is seeking a Training Infrastructure Engineer to design, build, and optimize infrastructure for large-scale model training. You will collaborate with AI researchers and engineers to create robust training pipelines, optimize distributed workloads, and ensure reliable model development.

Join a team scaling GPU clusters across data centers, implementing monitoring, logging, and automation, and advancing cost-efficient, high-performance training for multimodal and

Qualifications

  • Bachelor's degree in CS, CE, or related field, or equivalent practical experience.
  • 3+ years of experience with distributed systems and ML infrastructure.
  • Experience with PyTorch.
  • Proficiency in cloud platforms (AWS, GCP, Azure).
  • Experience with containerization, orchestration (Kubernetes, Docker).
  • Knowledge of distributed training techniques (data parallelism, model parallelism, FSDP).

Responsibilities

  • Design and implement scalable infrastructure for large-scale model training workloads.
  • Develop and maintain distributed training pipelines for LLMs and multimodal models.
  • Optimize training performance across multiple GPUs, nodes, and data centers.
  • Implement monitoring, logging, and debugging tools for training operations.
  • Architect and maintain data storage solutions for large-scale training datasets.
  • Automate infrastructure provisioning, scaling, and orchestration for model training.
  • Collaborate with researchers to implement and optimize training methodologies.
  • Analyze and improve efficiency, scalability, and cost-effectiveness of training systems.
  • Troubleshoot complex performance issues in distributed training environments.

Skills

Distributed systems
ML infrastructure
PyTorch
AWS
GCP
Azure
Kubernetes
Docker
Data parallelism
Model parallelism

Education

Bachelor's degree in Computer Science, Computer Engineering, or related field

Tools

Kubernetes
Docker

Job description

Fireworks AI in the United States is seeking a Training Infrastructure Engineer to design, build, and optimize infrastructure for large-scale model training. You will collaborate with AI researchers and engineers to create robust training pipelines, optimize distributed workloads, and ensure reliable model development.

Join a team scaling GPU clusters across data centers, implementing monitoring, logging, and automation, and advancing cost-efficient, high-performance training for multimodal and

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Research Engineer, AI Systems & Distributed Training
Research Engineer, AI Systems & Distributed Training

Fireworks • San Mateo (CA)

On-site
USD 140,000 - 200,000
AI Training Infrastructure Engineer - Scale LLM Training
AI Training Infrastructure Engineer - Scale LLM Training

Fireworks AI • San Mateo (CA)

On-site
USD 175,000 - 220,000
Equity
Comprehensive benefits
Competitive salary
Research Engineer: AI Training Infrastructure & Innovation
Research Engineer: AI Training Infrastructure & Innovation

Fireworks AI • New York (NY)

On-site
USD 150,000 - 230,000
Staff Research Engineer - Scalable ML Training Systems
Staff Research Engineer - Scalable ML Training Systems

Fireworks AI • San Mateo (CA)

On-site
USD 250,000 - 400,000
Member of Technical Staff, AI Training Infrastructure
Member of Technical Staff, AI Training Infrastructure

Fireworks AI • San Mateo (CA)

On-site
USD 175,000 - 220,000
Equity
Comprehensive benefits
Competitive salary
Software Engineer, AI Infrastructure & ML Systems
Software Engineer, AI Infrastructure & ML Systems

Fireworks AI • New York (NY)

On-site
USD 175,000 - 220,000
Equity
Competitive salary
Comprehensive benefits
AI Infrastructure Engineer: Scale & Serve Generative AI
AI Infrastructure Engineer: Scale & Serve Generative AI

Fireworks • San Mateo (CA)

On-site
USD 140,000 - 210,000
AI Infrastructure Engineer - Scalable ML Platform
AI Infrastructure Engineer - Scalable ML Platform

Fireworks AI • San Mateo (CA)

On-site
USD 175,000 - 220,000
Senior AI Training Infra Engineer - Scale GPU Clusters
Senior AI Training Infra Engineer - Scale GPU Clusters

Designworks Talent • Bellevue (WA)

Hybrid
USD 180,000 - 240,000
Medical insurance
401(k) with company match
Paid holidays
Staff Cloud Infrastructure Engineer, ML & AI Platforms
Staff Cloud Infrastructure Engineer, ML & AI Platforms

Fireworks AI • San Mateo (CA)

On-site
USD 175,000 - 220,000