Member of Technical Staff

Fireworks AI

United States

On-site

USD 140,000 - 220,000

Full time

8 days ago
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Job summary

Fireworks AI is hiring a Training Infrastructure Engineer to design, build, and maintain large-scale backend and cloud-native infrastructure for distributed ML training, inference, and data processing.

You will lead technical design discussions, mentor engineers, and drive best practices for scalable ML systems, focusing on reliability, low latency, and cost optimization. Proficiency with Docker, Kubernetes, and modern server-side languages is required.

Qualifications

  • Design, develop, and optimize large-scale backend infrastructure and distributed data systems in cloud environments.
  • Define and implement data-driven metrics to support company goals.
  • Lead cross-functional projects and contribute to technical design documentation.
  • Conduct coding interviews and provide structured feedback for engineering candidates.
  • Collaborate with ML, DevOps, and product teams to translate requirements into infrastructure solutions.
  • Experience with cloud-native tooling and infrastructure (Docker, Kubernetes).
  • Proficiency in major server-side languages (Python, C++, Go, TypeScript) and data processing/API systems (gRPC/Thrift).

Responsibilities

  • Design, build, and maintain scalable backend infrastructure for ML training, inference, and data processing pipelines.
  • Architect scalable, resilient systems emphasizing reliability and low latency.
  • Lead design discussions, mentor engineers, and establish best practices for ML infrastructure.
  • Optimize compute, storage, and network performance to reduce costs and improve efficiency.
  • Collaborate with ML, DevOps, and product teams to convert research requirements into robust infra.
  • Evaluate and integrate cloud-native/open-source tech (Kubernetes, Ray, Kubeflow, MLFlow).
  • Own end-to-end systems from design to deployment with fault tolerance and operations in mind.

Skills

Backend design
Distributed data systems
AB testing
Cross-functional collaboration
Mentoring engineers
Interviewing engineers

Education

Bachelor's degree in Computer Science or related field

Tools

Docker
Kubernetes
Python
Go
TypeScript
C++
PostgreSQL
MySQL
DynamoDB
Apache Spark
Apache Kafka
Apache Flink
Ray
Kubeflow
MLFlow
gRPC
Thrift

Job description

Responsibilities
  • As a Training Infrastructure Engineer, you’ll design, develop, and maintain large-scale backend and cloud-native infrastructure to support distributed machine learning training, inference, and data processing pipelines for our generative AI platform
  • You’ll architect scalable, resilient backend infrastructure, lead technical design discussions, mentor engineers, and establish best practices for large-scale machine learning systems
  • Architect and build scalable, resilient backend infrastructure to support distributed training, inference, and data processing pipelines
  • Lead technical design discussions, mentor engineers, and establish best practices for large-scale machine learning systems
  • Design and implement core backend services with a focus on efficiency and low latency
  • Drive infrastructure optimization initiatives for compute cost, storage lifecycle management, and network performance
  • Collaborate with machine learning, DevOps, and product teams to translate research and product requirements into robust infrastructure solutions
  • Evaluate and integrate cloud-native and open-source technologies such as Kubernetes, Ray, Kubeflow, and MLFlow to enhance platform reliability
  • Own end-to-end systems from design to deployment, emphasizing reliability, fault tolerance, and operational excellence
Requirements
  • 4 years of experience designing, building, and optimizing large-scale backend infrastructure and distributed data systems (e.g., PostgreSQL, MySQL, DynamoDB, Apache Spark, Apache Flink, Apache Kafka) in cloud environments (AWS, GCP, Azure, or equivalent), including cloud-native platforms, core infrastructure components, and optimization techniques (caching, indexing, sharding, replication, transactions, ACID)
  • 2 years of experience defining and implementing data-driven metrics to support company or team goals
  • 3 years of experience conducting A/B testing and scientific experimentation (e.g., Statsig, Meta Deltoid, Optimizely) to measure software impact
  • 4 years of experience writing technical design documentation, leading cross-functional projects, and collaborating with cross-functional teams to achieve business impact
  • 3 years of experience conducting coding interviews and providing systematic feedback for engineering candidates
  • Bachelor’s degree or equivalent in Computer Science or related field plus four (4) years of experience in software engineering or related role
  • 2 years of experience with cloud-native tools and infrastructure, such as Docker and Kubernetes
  • 4 years of experience with major server-side programming languages and frameworks (e.g., Python, C++, Go, TypeScript)
  • 3 years of experience developing and maintaining data processing and API systems, including client-server communication frameworks (e.g., gRPC, Thrift)
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Member of Technical Staff, Cloud Infrastructure
Member of Technical Staff, Cloud Infrastructure

Fireworks AI • San Mateo (CA)

On-site
USD 180,000 - 240,000
Member of Technical Staff (Cloud Infrastructure)
Member of Technical Staff (Cloud Infrastructure)

Fireworks AI • United States

On-site
USD 180,000 - 260,000
Member of Technical Staff - ML Infra
Member of Technical Staff - ML Infra

Kindredventures • San Francisco (CA)

On-site
USD 160,000 - 220,000
Senior Software Engineer
Senior Software Engineer

Vizient, Inc • Edina (MN)

On-site
USD 100,000 - 130,000
Member of Technical Staff
Member of Technical Staff

Harrison Clarke • San Francisco (CA)

On-site
USD 180,000 - 280,000
Senior ML Infra Engineer - Large-Scale Training & Pipelines
Senior ML Infra Engineer - Large-Scale Training & Pipelines

Kindredventures • San Francisco (CA)

On-site
USD 160,000 - 220,000
Lead Backend Engineer - AI Services
Lead Backend Engineer - AI Services

Harnham • Austin (TX)

Hybrid
USD 190,000 - 260,000
Tech Lead, Data & Inference Engineer
Tech Lead, Data & Inference Engineer

Catalyst Labs • Greenwich (CT)

Hybrid
USD 120,000 - 160,000
Tech Lead, Data & Inference Engineer
Tech Lead, Data & Inference Engineer

Catalyst Labs • Jacksonville (FL)

Hybrid
USD 130,000 - 160,000
Tech Lead, Data & Inference Engineer
Tech Lead, Data & Inference Engineer

Catalyst Labs • Massachusetts

Hybrid
USD 120,000 - 160,000