Member of Technical Staff (Reliability Engineering)

Fireworks AI

United States

On-site

USD 140,000 - 210,000

Full time

14 days+
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Job summary

Fireworks AI seeks a Reliability Engineer to own and drive production resiliency across GPU-accelerated inference and training systems. You will define SLOs, error budgets, and readiness criteria while leading incident response and postmortems.

Collaboration with cloud infra, ML infrastructure, and product teams is essential to ensure scalable, fail-fast systems. You will automate observability, implement standard tooling, and reduce toil with robust reliability practices across multi-region

Qualifications

  • 5+ years with Linux internals, system performance troubleshooting, and networking fundamentals.
  • Bachelor’s or Master’s in CS/CE or equivalent practical experience.
  • Experience with GPU and ML infrastructure exposure is a plus.
  • Strong emphasis on incident response and postmortems.

Responsibilities

  • Define reliability standards: SLOs, error budgets, production readiness criteria.
  • Own incident management, postmortems, and observability standards across services.
  • Automate failure testing, alerting, and runbooks to reduce toil.
  • Collaborate with cloud infra, AI systems, and product teams to ensure resiliency.

Skills

Linux internals
System performance troubleshooting
Networking fundamentals
Python
Go
C++
Rust
Kubernetes
Terraform
Docker
Prometheus
Grafana
OpenTelemetry
SRE practices
Incident response

Education

Bachelor's or Master's in Computer Science, Computer Engineering, or equivalent

Tools

Kubernetes
Terraform
Docker
Prometheus
Grafana
OpenTelemetry

Job description

  • Fireworks AI is one of the industry leaders in inference and training for open models. Open models are how the rest of the world gets to build on frontier AI without handing the keys to a single vendor, and our job is to make them fast, cheap, and dependable enough that this is a real choice
  • That work is systems work: GPU scheduling, kernel and runtime performance, networking, storage, Linux. We serve over 40 trillion tokens a day doing it
  • Reliability Engineering makes sure that platform runs dependably as it grows
  • You will work across cloud infrastructure, AI systems, and product teams to make sure the pieces fit together, fail gracefully, and hold up under load
  • You own the bar. You define what “reliable” means at Fireworks: SLOs, error budgets, production readiness, on-call expectations. Then you drive adoption across engineering
  • You own the process and the tooling. Incident management, postmortems, observability standards, failure testing, guardrails, and automation are yours end to end
  • Every team owns the reliability of what they build. You make that ownership practical. Structured logging and aggregation, metrics and tracing that work the same way everywhere, alerting that routes to the right owner, dashboards that answer “why is this slow.”
  • You choose where the leverage is. You have a wide view of the platform and the latitude to spend your time where it changes outcomes most
  • Define reliability standards: SLOs, error budgets, production readiness criteria. Not written in a vacuum: you instrument the systems and read the real telemetry the numbers come from
  • Own the reliability toolchain: Logging and telemetry pipelines, alerting standards, failure injection, load testing, self-healing automation, and AI-assisted investigation tooling
  • Keep customer experience from falling through the cracks: Per-service reliability is necessary but not sufficient. A customer can hit a bad experience while every system sits inside its SLO. You make sure those failures get an owner and a fix
  • Own the seams: The hardest failures live between systems: retries that amplify load, timeouts that do not compose, dependencies nobody mapped. You find them before customers do and drive fixes through the teams that own them
  • Run incident management: Coordinate live production issues, run blameless postmortems, and track follow-ups to completion
  • Reduce toil: Automate repetitive operational work so growth does not turn into an unsustainable on-call load
  • Partner across the org: Cloud infrastructure on capacity and multi-region risk, inference and training on failure modes in the serving and training stacks, performance on zero-downtime rollouts, product and control plane on customer-facing reliability
  • Solve Hard Problems: Tackle challenges at the forefront of AI infrastructure, from low-latency inference to scalable model serving
  • Build What’s Next: Work with bleeding-edge technology that impacts how businesses and developers harness AI globally
  • Cloud-native operations: Operating and debugging Kubernetes, Terraform, and Docker in high-throughput production
  • Reliability fundamentals: Fault-tolerant design, SLO/SLA management, automated failover, high-availability architecture
  • Breadth over comfort: Willingness to dig into unfamiliar parts of the stack when a problem crosses boundaries
  • Systems fundamentals: 5+ years with Linux internals, system performance troubleshooting, and networking fundamentals (TCP/IP, HTTP, gRPC)
  • Education: Bachelor’s or Master’s in Computer Science, Computer Engineering, or equivalent practical experience
  • Distributed systems: High-throughput control planes, microservices, or multi-region setups
  • Software engineering: 5+ years in Python, Go, C++, or Rust, writing production-grade tools and systems code
  • Influence without authority: You can get other teams to adopt a standard through credibility and useful tooling rather than mandate
  • Observability tooling: Prometheus, Grafana, OpenTelemetry, and alerting people actually act on
  • GPU and ML infrastructure exposure: GPUs, inference serving, or distributed training
  • AI-assisted operations: Building agents or LLM-based tooling for investigation, triage, or automation
  • Open source background: Contributions to infrastructure, systems, or ML serving projects
  • Startup agility: Comfortable where pragmatism and teamwork matter more than process
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Member of Technical Staff - Reliability Engineering
Member of Technical Staff - Reliability Engineering

Fireworks AI • New York (NY)

On-site
USD 140,000 - 210,000
Member of Technical Staff - Reliability Engineering
Member of Technical Staff - Reliability Engineering

Fireworks • San Mateo (CA)

On-site
USD 150,000 - 230,000
Senior Reliability Engineer – AI Infrastructure
Senior Reliability Engineer – AI Infrastructure

Fireworks AI • United States

On-site
USD 140,000 - 210,000
Senior Reliability Engineer — AI Infrastructure & SRE
Senior Reliability Engineer — AI Infrastructure & SRE

Fireworks • San Mateo (CA)

On-site
USD 150,000 - 230,000
Member of Technical Staff (Cloud Infrastructure)
Member of Technical Staff (Cloud Infrastructure)

Fireworks AI • United States

On-site
USD 180,000 - 260,000
IT DevOps Engineer
IT DevOps Engineer

Engg • San Mateo (CA)

On-site
USD 180,000 - 240,000
Member of Technical Staff, Enterprise Foundations
Member of Technical Staff, Enterprise Foundations

Fireworks AI • New York (NY)

On-site
USD 180,000 - 260,000
Member of Technical Staff
Member of Technical Staff

Fireworks AI • New York (NY)

On-site
USD 170,000 - 260,000
Member of Technical Staff (Enterprise Foundations)
Member of Technical Staff (Enterprise Foundations)

Fireworks AI • United States

On-site
USD 120,000 - 190,000
Software Engineer (LLM Infrastructure)
Software Engineer (LLM Infrastructure)

Fireworks AI • United States

On-site
USD 150,000 - 210,000