Senior Reliability Engineer — AI Infrastructure & SRE

Fireworks

San Mateo (CA)

On-site

USD 150,000 - 230,000

Full time

14 days+
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Job summary

Fireworks AI is seeking a Reliability Engineer to shape the platform’s dependability as it scales. You will own SLOs, error budgets, and production readiness, defining the reliability bar for services and guiding teams to meet it.

The role spans cloud infrastructure, AI systems, and product teams to ensure fast, cheap, and dependable model serving. Ideal candidates bring 5+ years in systems engineering, proficiency with Linux, Python, Go, C++, or Rust, and hands-on experience with Kubernetes,

Qualifications

  • 5+ years with Linux internals, system performance troubleshooting, and networking fundamentals (TCP/IP, HTTP, gRPC).
  • 5+ years in Python, Go, C++, or Rust, writing production-grade tools and systems code.
  • Operating and debugging Kubernetes, Terraform, and Docker in high-throughput production.
  • Distributed systems: High-throughput control planes, microservices, or multi-region setups.
  • Fault-tolerant design, SLO/SLA management, automated failover, high-availability architecture.
  • Influence without authority: You can get other teams to adopt a standard through credibility and useful tooling rather than mandate.
  • Breadth over comfort: Willingness to dig into unfamiliar parts of the stack when a problem crosses boundaries.

Responsibilities

  • Define reliability standards: SLOs, error budgets, production readiness criteria. Not written in a vacuum: you instrument the systems and read the real telemetry the numbers come from.
  • Own the reliability toolchain: Logging and telemetry pipelines, alerting standards, failure injection, load testing, self-healing automation, and AI-assisted investigation tooling.
  • Keep customer experience from falling through the cracks: Per-service reliability is necessary but not sufficient. A customer can hit a bad experience while every system sits inside its SLO. You make sure those failures get an owner and a fix.
  • Own the seams: The hardest failures live between systems: retries that amplify load, timeouts that do not compose, dependencies nobody mapped. You find them before customers do and drive fixes through the teams that own them.
  • Run incident management: Coordinate live production issues, run blameless postmortems, and track follow-ups to completion.
  • Reduce toil: Automate repetitive operational work so growth does not turn into an unsustainable on-call load.
  • Partner across the org: Cloud infrastructure on capacity and multi-region risk, inference and training on failure modes in the serving and training stacks, performance on zero-downtime rollouts, product and control plane on customer-facing reliability.

Skills

Linux internals
Performance troubleshooting
Networking fundamentals
Python
Go
C++
Rust
Kubernetes
Terraform
Docker
Distributed systems
SLO/SLA management
Fault-tolerant design
Influence without authority
Broad stack knowledge

Education

Bachelor’s or Master’s in Computer Science/Engineering

Tools

Prometheus
Grafana
OpenTelemetry

Job description

Fireworks AI is seeking a Reliability Engineer to shape the platform’s dependability as it scales. You will own SLOs, error budgets, and production readiness, defining the reliability bar for services and guiding teams to meet it.

The role spans cloud infrastructure, AI systems, and product teams to ensure fast, cheap, and dependable model serving. Ideal candidates bring 5+ years in systems engineering, proficiency with Linux, Python, Go, C++, or Rust, and hands-on experience with Kubernetes,

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Site Reliability Engineer — AI Platform Scale
Senior Site Reliability Engineer — AI Platform Scale

Future Secure AI • Austin (TX)

On-site
USD 140,000 - 190,000
Remote Senior Backend Engineer — AI-Driven Reliability
Remote Senior Backend Engineer — AI-Driven Reliability

Affirm • Richmond (VA)

On-site
USD 173,000 - 233,000
Health care coverage
Flexible Spending Wallets
Time off
+1
Senior Staff SRE Engineer: Scale Reliability & AI Ops
Senior Staff SRE Engineer: Scale Reliability & AI Ops

Wand AI • Palo Alto (CA)

On-site
USD 130,000 - 180,000
AI Reliability Engineer (AI SRE)
AI Reliability Engineer (AI SRE)

DeWinter Group • Campbell (CA)

Remote
Senior Backend Reliability Engineer — AI‑Driven Platform (Remote)
Senior Backend Reliability Engineer — AI‑Driven Platform (Remote)

Affirm • Riverside (OH)

Remote
USD 173,000 - 233,000
Health coverage
FSA Wallets
Time off
+1
Platform Reliability Engineer – Scalable AI Infra
Platform Reliability Engineer – Scalable AI Infra

OpenAI • California (MO)

On-site
USD 190,000 - 270,000
Relocation assistance
Member of Technical Staff - Reliability Engineering
Member of Technical Staff - Reliability Engineering

Fireworks • San Mateo (CA)

On-site
USD 150,000 - 230,000
Senior Backend Engineer AI-Driven Reliability (Remote)
Senior Backend Engineer AI-Driven Reliability (Remote)

Affirm • Boise (ID)

On-site
USD 173,000 - 233,000
Health coverage for you and dependents
FSAs - tech spending
Time off
+1
Senior SRE: AI-Driven Infra & Reliability
Senior SRE: AI-Driven Infra & Reliability

Jobless • Ann Arbor (MI)

Hybrid
USD 180,000 - 240,000
Health Care Coverage
Life Insurance
Health Savings Account
+3
Senior Backend Engineer - AI-Driven Reliability Platform
Senior Backend Engineer - AI-Driven Reliability Platform

Affirm • Austin (TX)

On-site
USD 173,000 - 233,000
Health care coverage
Flexible Spending Wallets
Time off
+1