Founding Engineer - Site Reliability

uRun

San Francisco (CA)

On-site

USD 120,000 - 140,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Health, dental, and vision – full coverage
401(k) – company-supported retirement savings
FSA/HSA – flexible spending accounts
Paid time off
Top-tier tooling
MacBook Pro and AirPods

Job summary

uRun is seeking a Site Reliability Engineer to own the reliability culture from scratch. As a founding hire, you'll define the observability stack and incident response processes, working closely with infrastructure engineers to ensure uptime and efficiency.

We're building cutting-edge infrastructure for interactive AI, and this role is pivotal in shaping our foundation. A strong background in SRE, cloud infrastructure, and incident management is essential.

Enjoy competitive salary and equity, along with comprehensive health coverage, flexible spending accounts, and the best tools available.

Qualifications

  • 7+ years in site reliability, production engineering, or infrastructure engineering in a high-availability, low-latency environment.
  • Deep experience owning SLOs, error budgets, and on-call processes in production at scale.
  • Hands-on with Kubernetes and cloud infrastructure; able to debug issues effectively.

Responsibilities

  • Define and own SLOs and error budgets across uRun's inference platform.
  • Build and maintain the observability stack end-to-end.
  • Lead incident response from detection to resolution and postmortems.

Skills

Site reliability engineering
Production engineering
Infrastructure engineering
Observability stacks (Prometheus, Grafana, Datadog)
Kubernetes
Cloud infrastructure (AWS preferred)
Incident response

Job description

The problem we saw

Most AI infrastructure is built for batch: send a query, wait, get a response, reset. Powerful, but transactional. AI is becoming interactive — sessions that hold state, models that stay alive between turns, generation that responds as it runs — and the infrastructure to deliver that at scale doesn't really exist yet.

The bottleneck isn't the models anymore. It's the infrastructure underneath them.

What we're building to fix it

uRun is the inference cloud for interactive AI: the compute layer that makes real-time, stateful inference possible at scale. We came out of stealth in April 2026, are backed by top-tier investors, and are founded by Keegan McCallum, who scaled inference infrastructure for some of the most demanding generative AI workloads in production.

We're an infrastructure company. We build the layer that model labs, builders, and research teams ship on top of.

Where you come in

Reliability at uRun isn't a feature — it's the product. When model labs and production teams build on top of our inference platform, they are trusting us with their uptime, their latency, and their users. As our Site Reliability Engineer, you will own that trust end-to-end.

This is a founding SRE hire. You will define the reliability culture from scratch: the observability stack, the incident response playbooks, the SLOs, and the on-call process. You will work directly with infrastructure and platform engineers to close the gap between what we ship and what stays up.

What you’ll actually be doing day-to-day
  • Define and own SLOs and error budgets across uRun’s inference platform and supporting infrastructure
  • Build and maintain the observability stack end-to-end: metrics, logging, tracing, and alerting across a distributed GPU compute environment
  • Lead incident response: detection, triage, resolution, and blameless postmortems that drive lasting fixes
  • Partner with ML infrastructure engineers to embed reliability into the deployment pipeline from day one
  • Design and maintain runbooks, on-call rotations, and escalation paths as the team scales
  • Drive capacity planning and traffic management across heterogeneous compute to protect latency and availability under load
  • Identify and eliminate toil through automation, building systems that scale without scaling the team proportionally
What skills you need for the journey
  • 7+ years in site reliability, production engineering, or infrastructure engineering in a high-availability, low-latency environment
  • Deep experience owning SLOs, error budgets, and on-call processes in production at scale
  • Strong observability background: you have built or owned monitoring stacks (Prometheus, Grafana, Datadog, or equivalent) and know what good alerting looks like
  • Proven incident response experience: you have led real incidents under pressure and written postmortems that actually changed behaviour
  • Hands‑on with Kubernetes and cloud infrastructure (AWS preferred): you can debug a failing pod and a misconfigured VPC in the same afternoon
  • Strong software engineering fundamentals: you write automation, not just runbooks
  • Comfortable operating as the first and only SRE, setting standards without a template to follow
Things that will give you an edge
  • Experience supporting GPU compute or ML inference infrastructure in production
  • Familiarity with stateful workloads, long‑running sessions, or streaming inference systems
  • Exposure to multi‑tenant platforms where isolation, noisy neighbour problems, and billing‑aware scheduling matter
  • Prior founding or sole SRE experience at an early‑stage company
What you’ll get in return

Competitive salary and meaningful equity in an early‑stage AI infrastructure company. The band above is our target; for an exceptional candidate we’ll go higher. Equity is real — you’re early, and the grant reflects that.

  • Health, dental, and vision — full coverage
  • 401(k) — company‑supported retirement savings
  • FSA/HSA — flexible spending accounts for healthcare costs
  • Paid time off — we trust you to manage your time
  • Top‑tier tooling — access to the best AI tools available: Claude, Codex, Kimi, and whatever else helps you move faster
  • MacBook Pro and AirPods — the hardware you need, on us
How we work (and what that feels like day-to-day)

We build the stage, not the show. We’re an infrastructure company, a developer‑tools company, and a production partner for model labs, and focus is a deliberate choice we’ve made and hold to.

Day‑to‑day, that means a small team, a high bar, and real ownership. You won’t wait for permission or inherit a backlog of someone else’s decisions; in a founding security role, the function is what you make it.

It also means ambiguity: priorities shift, not everything is documented, and you’ll often be the person who decides what "secure enough, for now" means. That suits some people and not others, and we’d rather you know that before you apply.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Founding Engineer - Platform
Founding Engineer - Platform

uRun • San Francisco (CA)

On-site
USD 140,000 - 180,000
Competitive salary and equity
Full health, dental, and vision coverage
401(k) retirement savings
+4
Founding Engineer - ML Performance
Founding Engineer - ML Performance

uRun • San Francisco (CA)

On-site
USD 120,000 - 160,000
Health, dental, and vision
401(k) participation
Flexible spending accounts
+3
Senior Site Reliability Engineer
Senior Site Reliability Engineer

hardrockdigital • United States

Hybrid
USD 120,000 - 160,000
Competitive pay and benefits
Flexible vacation allowance
Startup culture with global brand support
+1
Site Reliability Engineer
Site Reliability Engineer

Runloop • San Francisco (CA)

Hybrid
USD 150,000 - 180,000
Health insurance
Daily catered lunch
Equity & competitive salary
Founding SRE for Interactive AI Inference Infra
Founding SRE for Interactive AI Inference Infra

uRun • San Francisco (CA)

On-site
USD 120,000 - 140,000
Health, dental, and vision – full coverage
401(k) – company-supported retirement savings
FSA/HSA – flexible spending accounts
+3
Head of SRE
Head of SRE

Wand AI • Palo Alto (CA)

On-site
USD 130,000 - 180,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Juniper Square • United States

Remote
USD 165,000 - 195,000
Health, dental, and vision care
Life insurance
Mental wellness coverage
+4
Site Reliability Engineer - NYC
Site Reliability Engineer - NYC

Mistral • New York (NY)

Hybrid
USD 140,000 - 190,000
Competitive salary and equity
Healthcare: Medical/Dental/Vision for你
401K with match
+6
Staff Site Reliability Engineer – Automation and Platform
Staff Site Reliability Engineer – Automation and Platform

Cerebras • Sunnyvale (CA)

On-site
USD 150,000 - 200,000
Senior Site Reliability Engineer II
Senior Site Reliability Engineer II

Juniper Square • United States

On-site
USD 165,000 - 195,000
Health, dental, and vision care
Life insurance
Mental wellness coverage
+3