Senior Reliability Engineer, AI Infrastructure

Fireworks

San Mateo

On-site

PHP 900,000 - 1,300,000

Full time

8 days ago

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Fireworks is seeking a Reliability Engineer to ensure the platform runs reliably as it scales. You will define SLOs, own the reliability toolchain, and coordinate incident response across cloud infra, AI systems, and product teams.

You’ll tackle failures that span multiple services, drive automation to reduce toil, and work with cross-functional teams to uphold production readiness and high availability. A strong foundation in distributed systems and modern open models is essential.

Qualifications

  • 5+ years with Linux internals, system performance troubleshooting, and networking fundamentals (TCP/IP, HTTP, gRPC).
  • 5+ years in Python, Go, C++, or Rust, writing production-grade tools and systems code.
  • Cloud-native operations: Kubernetes, Terraform, and Docker in high-throughput production.
  • Distributed systems: High-throughput control planes, microservices, or multi-region setups.
  • Reliability fundamentals: Fault-tolerant design, SLO/SLA management, automated failover, HA architecture.
  • Education: Bachelor's or Master's in CS/CE or equivalent practical experience.

Responsibilities

  • Define reliability standards: SLOs, error budgets, production readiness criteria, instrumentation of systems.
  • Own the reliability toolchain: Logging and telemetry pipelines, alerting standards, failure injection, load testing, automation.
  • Keep customer experience from falling through the cracks: Ensure per-service reliability aligns with overall SLOs.
  • Own the seams: Identify cross-system failures, manage retries, timeouts, and dependencies.
  • Run incident management: Coordinate live production issues and postmortems.
  • Reduce toil: Automate repetitive operational work.
  • Partner across the org: Align cloud infra, AI serving, and product control planes on reliability.

Skills

Linux internals
Python
Go
C++
Rust
Networking fundamentals
SRE practices

Education

Bachelor's or Master's in CS/CE or equivalent

Tools

Kubernetes
Docker
Terraform
OpenTelemetry
Prometheus

Job description

Fireworks is seeking a Reliability Engineer to ensure the platform runs reliably as it scales. You will define SLOs, own the reliability toolchain, and coordinate incident response across cloud infra, AI systems, and product teams.

You’ll tackle failures that span multiple services, drive automation to reduce toil, and work with cross-functional teams to uphold production readiness and high availability. A strong foundation in distributed systems and modern open models is essential.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Backend Engineer, AI Platform & Orchestration
Senior Backend Engineer, AI Platform & Orchestration

Fireworks • San Mateo

On-site
PHP 900,000 - 1,300,000
Member of Technical Staff - Reliability Engineering
Member of Technical Staff - Reliability Engineering

Fireworks • San Mateo

On-site
PHP 900,000 - 1,300,000
Senior AI-Driven Cloud Reliability Engineer
Senior AI-Driven Cloud Reliability Engineer

IgniteTech • Philippines

On-site
PHP 1,800,000 - 3,600,000
AI Infrastructure Security Operations Lead
AI Infrastructure Security Operations Lead

Fireworks • San Mateo

On-site
PHP 1,800,000 - 3,200,000
Member of Technical Staff, Software Engineer
Member of Technical Staff, Software Engineer

Fireworks • San Mateo

On-site
PHP 900,000 - 1,300,000
Senior Data Platform Reliability Engineer
Senior Data Platform Reliability Engineer

TerraBarn Inc. • Cebu City

On-site
PHP 1,200,000 - 2,400,000
HMO
Security Operations Lead
Security Operations Lead

Fireworks • San Mateo

On-site
PHP 1,800,000 - 3,200,000
Lead Site Reliability Engineer (SRE)
Lead Site Reliability Engineer (SRE)

EPAM Systems • Mexico

On-site
PHP 5,846,000 - 8,616,000
SR. Data Platform Reliability Engineer/Data SRE (Permanent)- Onsite
SR. Data Platform Reliability Engineer/Data SRE (Permanent)- Onsite

ATS CONSULTING SERVICES PH INC. • Mandaluyong

On-site
PHP 1,000,000 - 2,000,000
Data Platform Reliability Engineer - Kubernetes & Cloud
Data Platform Reliability Engineer - Kubernetes & Cloud

TerraBarn Inc. • Mandaluyong

On-site
PHP 1,200,000 - 2,400,000
HMO