Site Reliability Engineer — Ultra-Low Latency AI Infra

Blaxel (YC X25)

San Francisco (CA)

On-site

USD 175,000 - 250,000

Full time

14 days+
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Job summary

Blaxel (YC X25) is seeking a Site Reliability Engineer in San Francisco, California, to ensure the reliability of our AI infrastructure platform. Your mission is to operate the core systems that power ultra-low-latency serverless computing as we serve billions of requests.

With a focus on automating operations, you will design systems for monitoring and incident response, aiming for world-class reliability. Join us to push the boundaries of AI technology.

Qualifications

  • 3+ years in SRE, DevOps, or infrastructure engineering.
  • Strong proficiency in Go, Rust, or Python.
  • Hands-on experience with a major cloud provider.
  • Solid knowledge of Linux systems and networking.

Responsibilities

  • Architect, operate, and improve the core infrastructure.
  • Build and evolve the observability stack.
  • Define, monitor, and drive SLOs/SLIs.
  • Lead incident response with rigor.
  • Design self-healing operational systems.

Skills

SRE
DevOps
Python
Go
Rust
Kubernetes

Tools

AWS
Terraform
Linux

Job description

Blaxel (YC X25) is seeking a Site Reliability Engineer in San Francisco, California, to ensure the reliability of our AI infrastructure platform. Your mission is to operate the core systems that power ultra-low-latency serverless computing as we serve billions of requests.

With a focus on automating operations, you will design systems for monitoring and incident response, aiming for world-class reliability. Join us to push the boundaries of AI technology.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Site Reliability Engineer
Site Reliability Engineer

Blaxel (YC X25) • San Francisco (CA)

On-site
USD 175,000 - 250,000
Senior Site Reliability Engineer — AI Platform Scale
Senior Site Reliability Engineer — AI Platform Scale

Future Secure AI • Austin (TX)

On-site
USD 140,000 - 190,000
Site Reliability Engineer — ML Infra, Scale & Equity
Site Reliability Engineer — ML Infra, Scale & Equity

Baseten • New York (NY)

On-site
USD 165,000 - 330,000
Site Reliability Engineer — Scale & Resilience for AI Ops
Site Reliability Engineer — Scale & Resilience for AI Ops

HappyRobot • San Francisco (CA)

On-site
USD 120,000 - 160,000
Site Reliability Engineer
Site Reliability Engineer

FLUIX • Palo Alto (CA)

On-site
USD 120,000 - 150,000
Attractive compensation package including equity options
Comprehensive health, dental, and vision insurance
Opportunities for professional growth
Site Reliability Engineer – AI Platform Infra (Onsite NYC)
Site Reliability Engineer – AI Platform Infra (Onsite NYC)

getbasis.ai • New York (NY)

On-site
USD 140,000 - 190,000
Health & Wellness benefits
Time off — unlimited PTO + holidays
In-Office perks — meals, kitchen, desk
Site Reliability Engineer — AI Inference at Scale
Site Reliability Engineer — AI Inference at Scale

Linuxconfig • San Francisco (CA), Northern (KY)

Hybrid
USD 200,000 - 400,000
Health benefits
Dental benefits
Vision benefits
+1
Senior SRE - AI Platform Reliability & Scale
Senior SRE - AI Platform Reliability & Scale

Blitzy • Cambridge (MA)

On-site
USD 120,000 - 150,000
Ownership and equity
Dynamic work environment
Growth opportunities
Senior ML Infra Engineer — Real-Time, Low-Latency
Senior ML Infra Engineer — Real-Time, Low-Latency

LMArena • San Francisco (CA)

On-site
USD 170,000 - 260,000
Competitive compensation
Comprehensive health benefits
Opportunity to work on cutting-edge AI
Director, AI Infrastructure & Server Systems
Director, AI Infrastructure & Server Systems

Axelera AI • Town of Norway (WI)

On-site
USD 140,000 - 190,000
Pension plan
Employee insurances
Option for company shares