Staff Reliability Engineer for AI Infrastructure

Socket.dev

San Mateo (CA)

On-site

USD 180,000 - 240,000

Full time

7 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Fireworks AI is seeking a Reliability Engineer to own the bar for platform reliability across cloud infrastructure, AI systems, and product teams. You will define SLOs, lead incident response, and drive improvements that keep services fast, available, and dependable across the stack.

You will work across Linux, container platforms, and distributed systems, building automation and observability that scales with growth.

Qualifications

  • 5+ years Linux internals, performance troubleshooting, and networking fundamentals.
  • 5+ years in Python, Go, C++, or Rust, building production-grade tools.
  • Experience with Kubernetes, Terraform, and Docker in high-throughput production.
  • Distributed systems focus with fault-tolerant design and SLO/SLA management.
  • Bachelor's or Master's in CS, CS Engineering, or equivalent practical experience.

Responsibilities

  • Define reliability standards: SLOs, error budgets, production readiness criteria.
  • Own the reliability toolchain: Logging and telemetry pipelines, alerting, failure injection, load testing, automation.
  • Keep customer experience by ensuring per-service reliability and accountability.
  • Own seams between systems; identify and fix cross-service failure modes.
  • Run incident management: coordinate live issues, postmortems, and follow-ups.
  • Reduce toil through automation and scalable observability.
  • Partner across org on capacity, multi-region risk, and serving/training stacks.

Skills

Linux internals
Performance troubleshooting
Networking fundamentals
Python
Go
C++
Rust
Kubernetes
Terraform
Docker
Distributed systems
SRE / Reliability
Education in CS

Education

Bachelor's or Master's in Computer Science or related field

Tools

Prometheus
Grafana
OpenTelemetry
LLM-based tooling
Git

Job description

Fireworks AI is seeking a Reliability Engineer to own the bar for platform reliability across cloud infrastructure, AI systems, and product teams. You will define SLOs, lead incident response, and drive improvements that keep services fast, available, and dependable across the stack.

You will work across Linux, container platforms, and distributed systems, building automation and observability that scales with growth.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Member of Technical Staff - Reliability Engineering
Member of Technical Staff - Reliability Engineering

Socket.dev • San Mateo (CA)

On-site
USD 180,000 - 240,000
Software Engineer, AI Infrastructure & ML Systems
Software Engineer, AI Infrastructure & ML Systems

Fireworks AI • New York (NY)

On-site
USD 175,000 - 220,000
Equity
Competitive salary
Comprehensive benefits
Staff Cloud Infrastructure Engineer, ML & AI Platforms
Staff Cloud Infrastructure Engineer, ML & AI Platforms

Fireworks AI • San Mateo (CA)

On-site
USD 175,000 - 220,000
AI Infrastructure Engineer - Scalable ML Platform
AI Infrastructure Engineer - Scalable ML Platform

Fireworks AI • San Mateo (CA)

On-site
USD 175,000 - 220,000
Staff Cloud Infrastructure Engineer — ML & Distributed Systems
Staff Cloud Infrastructure Engineer — ML & Distributed Systems

Fireworks AI • New York (NY)

On-site
USD 175,000 - 220,000
Senior AI-Driven Infra & Reliability Engineer – Remote
Senior AI-Driven Infra & Reliability Engineer – Remote

IgniteTech • Germany (OH)

On-site
USD 130,000 - 170,000
AI-first culture
No limits on AI tooling or compute
Fully remote
+2
Senior Backend Engineer AI-Driven Reliability (Remote)
Senior Backend Engineer AI-Driven Reliability (Remote)

Affirm • Boise (ID)

On-site
USD 173,000 - 233,000
Health coverage for you and dependents
FSAs - tech spending
Time off
+1
Chief AI-Driven SRE & Reliability Leader
Chief AI-Driven SRE & Reliability Leader

IgniteTech • United States

Remote
USD 150,000 - 200,000
Fully remote work
No tooling limits
Remote Senior Backend Engineer — AI-Driven Reliability
Remote Senior Backend Engineer — AI-Driven Reliability

Affirm • Richmond (VA)

On-site
USD 173,000 - 233,000
Health care coverage
Flexible Spending Wallets
Time off
+1
Senior Backend Engineer - AI-Driven Reliability Platform
Senior Backend Engineer - AI-Driven Reliability Platform

Affirm • Miami (FL)

On-site
USD 173,000 - 233,000
Health and wellness benefits
Remote-first culture
Competitive equity