Provider Ops SRE: Reliability, Failover & Automation

OpenRouter, Inc

United States

Vor Ort

USD 150.000 - 210.000

Vollzeit

Vor 2 Tagen
Sei unter den ersten Bewerbenden
Bewerbungsgenerator

Eine komplette Bewerbung in einer Minute — maßgeschneiderter Lebenslauf und Anschreiben, fertig zum Versenden.

Schaffe es an den ATS-Filtern vorbei

Zusammenfassung

OpenRouter, Inc. is hiring an AI Inference SRE to own the operational health of our provider supply. You’ll build monitoring, run on-call, and improve failover across 80+ providers and thousands of endpoints.

You’ll join the Provider Operations team and report to the Provider Operations Manager. Candidates should have strong observability experience and a product mindset for reliability in high-traffic AI workloads.

Qualifikationen

  • 4+ years in SRE or production engineering.
  • Observability with metrics, tracing, logs and SLOs.
  • Software engineer who writes tools (TypeScript/Python).
  • Experience with distributed systems failure modes.
  • Calm incident commander who communicates under pressure.
  • Understanding of how LLM inference is served.

Aufgaben

  • Provider health and observability: monitor latency, throughput, error rates, uptime, and output correctness.
  • Detection and failover: quick degradation detection and traffic rerouting.
  • Incident response: on-call, triage, mitigations, postmortems.
  • Provider accountability: telemetry to scorecards and escalation point.
  • Quality regression detection: canaries and evals for regressions.
  • Automate the toil: replace manual provider-ops work with tooling.
  • Capacity and launch readiness: load-test endpoints before big launches.

Kenntnisse

Observability tooling
SRE / Production engineering
TypeScript
Python
Distributed systems
Incident response
LLM inference

Tools

Postgres
ClickHouse
GCP
Cloudflare Workers
Vercel

Jobbeschreibung

OpenRouter, Inc. is hiring an AI Inference SRE to own the operational health of our provider supply. You’ll build monitoring, run on-call, and improve failover across 80+ providers and thousands of endpoints.

You’ll join the Provider Operations team and report to the Provider Operations Manager. Candidates should have strong observability experience and a product mindset for reliability in high-traffic AI workloads.

Hol dir deinen kostenlosen, vertraulichen Lebenslauf-Check.

oder ziehe deine Datei hierhin.

Similar jobs

Ähnliche Jobs, die dir auch gefallen könnten

SRE, Provider Operations — AI Inference Reliability
SRE, Provider Operations — AI Inference Reliability

AI Chopping Block, Inc. • Northern (KY)

Hybrid
USD 120.000 - 160.000
SRE for AI Provider Operations: Resilient Routing
SRE for AI Provider Operations: Resilient Routing

OpenRouter, Inc • New York (NY)

Vor Ort
USD 140.000 - 190.000
Site Reliability Engineer, Provider Operations
Site Reliability Engineer, Provider Operations

OpenRouter, Inc • USA

Remote
USD 150.000 - 210.000
Site Reliability Engineer, Provider Operations
Site Reliability Engineer, Provider Operations

AI Chopping Block, Inc. • Northern (KY)

Hybrid
USD 120.000 - 160.000
Site Reliability Engineer, Provider Operations
Site Reliability Engineer, Provider Operations

OpenRouter, Inc • New York (NY)

Vor Ort
USD 140.000 - 190.000
Senior SRE Leader: AI-Driven Infra & Platform Reliability
Senior SRE Leader: AI-Driven Infra & Platform Reliability

Invoca • Los Angeles (CA)

Remote
USD 190.000 - 250.000
Health benefits
Mental wellbeing
Wellness subsidy
+6
AI Provider Operations & Support Engineer
AI Provider Operations & Support Engineer

OpenRouter • USA

Remote
USD 110.000 - 150.000
Senior SRE: AI Compute, Kubernetes & Observability
Senior SRE: AI Compute, Kubernetes & Observability

itMatch • USA

Remote
USD 140.000 - 190.000
Health benefits
Financial planning
Family benefits
+2
Senior SRE Leader - Remote-First Infra & AI Ops
Senior SRE Leader - Remote-First Infra & AI Ops

Invoca • Chicago (IL)

Remote
USD 190.000 - 250.000
Health benefits day-one
Wellness subsidy
Flexible time off
+5
SRE with AI (Don’t share any AI/ML profiles) – C2C jobs
SRE with AI (Don’t share any AI/ML profiles) – C2C jobs

Tech Mirrors • Mountain View (CA)

Vor Ort
USD 180.000 - 240.000