Site Reliability Engineer, Provider Operations

OpenRouter, Inc

New York (NY)

On-site

USD 140,000 - 190,000

Full time

36 hours ago
Be an early applicant
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Job summary

OpenRouter is hiring its first AI Inference SRE to own the operational health of provider routes. You’ll ensure every endpoint is fast, correct, and available across 80+ providers, with telemetry, SLOs, and alerting that prevent customers from noticing issues.

You’ll lead on-call incident response, automate toil, and build tooling to detect degradations early. The role combines software engineering (TypeScript/Python) with understanding of LLM serving, streaming, and provider APIs to keep

Qualifications

  • 4+ years in SRE, production engineering, or infra roles handling high-traffic systems.
  • Strong observability: metrics, tracing, logs, SLOs and actionable alerts.
  • Software engineer who writes tools, prefers TS/Python.
  • Experience with distributed failure modes: timeouts, retries, backpressure.
  • Calm incident commander who communicates under pressure.
  • Understanding of how LLM inference is served: streaming, prompt caching, throughput tradeoffs.

Responsibilities

  • Provider health and observability: build monitoring for latency, throughput, error rates, uptime, and correctness.
  • Detection and failover: improve detection of degraded endpoints and automatic traffic shifts.
  • Incident response: own on-call, triage, mitigate, postmortems, and follow-ups.
  • Provider accountability: turn telemetry into scorecards and SLO reporting.
  • Quality regression detection: build canaries and evals to catch regressions.
  • Automate the toil: replace manual provider-ops with safe tooling.
  • Capacity and launch readiness: load-test endpoints before big launches.

Skills

SRE / production engineering
Observability
Typescript
Python
Distributed systems
Incident command
LLM inference serving

Tools

Cloudflare Workers
Postgres
ClickHouse
GCP
Vercel

Job description

About OpenRouter

OpenRouter is the leading AI routing and infrastructure layer that enterprises use to access, manage, and optimize the best large language models across providers—without lock-in, capacity constraints, or unnecessary cost. We power the most advanced AI teams in the world by giving them the flexibility to move fast, scale confidently, and stay future-proof as models evolve.

As enterprise adoption of AI accelerates, OpenRouter sits at the center of how organizations operationalize LLMs across research, product, and production workloads.

About the Role

OpenRouter routes almost a billion requests and more than 20 trillion tokens a day, across 80+ providers and thousands of endpoints. Every one of those providers can degrade, rate-limit, change behavior, or go down without warning. Our customers count on us to absorb that chaos so their apps never notice.

We're hiring our first AI Inference SRE to own the operational health of our provider supply. You'll make sure every endpoint we route to is fast, correct, and available, and that we detect and route around problems before customers do. You'll sit on the Provider Operations team, reporting to the Provider Operations Manager.

What You'll Do
  • Provider health and observability. Build and own monitoring for every provider and endpoint: latency, throughput, error rates, uptime, and output correctness. Set SLOs per provider tier and alert on them.

  • Detection and failover. Improve how quickly we detect degraded endpoints, and work with the routing team so traffic shifts away from them automatically.

  • Incident response. Own on-call for provider incidents: triage, mitigate, communicate with providers, run postmortems, and drive follow-ups to closure.

  • Provider accountability. Turn telemetry into scorecards and SLO reporting that providers act on, and be the technical escalation point when a provider's endpoint is misbehaving.

  • Quality regression detection. Build continuous canaries and evals that catch silent regressions (quantization changes, broken tool calling, truncated streams, pricing or usage-reporting mismatches), not just outright downtime.

  • Automate the toil. Replace manual provider-ops work (disabling endpoints, capacity changes, deprecations, rate-limit tuning) with safe, auditable tooling.

  • Capacity and launch readiness. Build tooling to load-test endpoints before big launches so day-zero traffic doesn't take them down.

About You
  • 4+ years in SRE, production engineering, or infrastructure roles running high-traffic, customer-facing systems.

  • Strong with observability tooling and practice: metrics, tracing, logs, SLOs/error budgets, alerting that is always actionable.

  • Capable software engineer who prefers writing tools over executing runbooks. TypeScript and/or Python.

  • Experienced with distributed systems failure modes: timeouts, retries, backpressure, partial outages, noisy neighbors.

  • Calm, clear incident commander who communicates well with external partners under pressure.

  • Understands, or is eager to learn deeply, how LLM inference is served: streaming, tool calling, prompt caching, throughput/latency tradeoffs, and how provider APIs differ.

Nice to Have
  • Experience at an inference provider, model lab, GPU cloud, or API gateway/CDN company.

  • Experience with our stack: TypeScript, Cloudflare Workers, Postgres, ClickHouse, GCP, Vercel.

  • Background in routing, load balancing, or traffic management systems.

  • Experience with evals or synthetic monitoring for ML systems.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Site Reliability Engineer, Provider Operations
Site Reliability Engineer, Provider Operations

OpenRouter, Inc • United States

Remote
USD 150,000 - 210,000
Site Reliability Engineer, Provider Operations
Site Reliability Engineer, Provider Operations

AI Chopping Block, Inc. • Northern (KY)

Hybrid
USD 120,000 - 160,000
Engineering Manager, Provider Ecosystem Remote (US)
Engineering Manager, Provider Ecosystem Remote (US)

OpenRouter, Inc • Northern (KY)

Hybrid
USD 180,000 - 250,000
Restaurant example not provided
Engineering Manager, Provider Ecosystem
Engineering Manager, Provider Ecosystem

OpenRouter, Inc • New York (NY)

On-site
USD 180,000 - 260,000
AI Provider Operations & Support Engineer
AI Provider Operations & Support Engineer

OpenRouter • United States

Remote
USD 110,000 - 150,000
Software Engineer, Platform
Software Engineer, Platform

OpenRouter • United States

On-site
USD 215,000 - 285,000
Provider Ops SRE: Reliability, Failover & Automation
Provider Ops SRE: Reliability, Failover & Automation

OpenRouter, Inc • United States

On-site
USD 150,000 - 210,000
SRE for AI Provider Operations: Resilient Routing
SRE for AI Provider Operations: Resilient Routing

OpenRouter, Inc • New York (NY)

On-site
USD 140,000 - 190,000
SRE, Provider Operations — AI Inference Reliability
SRE, Provider Operations — AI Inference Reliability

AI Chopping Block, Inc. • Northern (KY)

Hybrid
USD 120,000 - 160,000
Software Engineer, Platform
Software Engineer, Platform

OpenRouter • New York (NY)

On-site
USD 215,000 - 285,000