Founding Engineer, Real-Time AI Runtime

uRun

United States

Remote

USD 180,000 - 260,000

Full time

14 days+
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Benefits offered by this job

Equity
Health, dental, and vision
401(k)
FSA/HSA
Paid time off
Top-tier tooling
MacBook Pro and AirPods

Job summary

uRun is building a universal runtime for real-time, stateful AI inference. We seek a senior engineer to design and own the scalable, low-latency infrastructure powering live, multi-user AI workloads. You’ll drive latency, memory usage, and QoS as a hands-on technical leader reporting to the founder.

You’ll shape platform standards, mentor engineers, and partner with ML and video teams to push the boundaries of real-time AI infrastructure in a fast-paced startup.

Qualifications

  • 7+ years of engineering with a track record architecting and owning large-scale production systems.
  • Deep Kubernetes expertise, including GPU-heavy clusters and service-mesh patterns.
  • Strong cloud and infrastructure-as-code: AWS, GCP, or Azure; Terraform, Pulumi; networking and security (VPC, IAM, API-gateway-style routing).
  • SRE-style thinking and observability depth: Prometheus/Grafana, OpenTelemetry, tracing, SLOs, incident response, post-mortems.
  • Proficiency in at least one of Python, Go, or TypeScript/Node.js for platform tooling, automation, and glue code.
  • Experience with streaming or real-time systems: WebRTC, low-latency video pipelines.
  • Track record mentoring engineers and influencing cross-functional teams.

Responsibilities

  • Design, operate, and evolve the cloud-native platform that runs uRun’s real-time inference and video runtime, Kubernetes, GPU-heavy workloads, and streaming pipelines.
  • Own observability, reliability, and performance at scale: SLO-driven capacity, autoscaling, failover, and cost-efficient GPU provisioning.
  • Build and maintain the platform primitives that product and ML teams depend on, service meshes, deployment pipelines, secrets and credential management, and configuration-as-code.
  • Partner closely with ML and video-workload engineers to optimise for low-latency inference, memory-bound workloads, and streaming data flows.
  • Define and champion platform standards for security, observability, and incident response, drawing on SRE-style practices.
  • Mentor and unblock other engineers, and act as a technical leader on architecture, trade-offs, and long-term platform evolution.

Skills

Kubernetes expertise
GPU workloads
Cloud infrastructure as code
SRE and observability
Python/Go/TypeScript
Streaming/real-time systems
Mentoring engineers

Tools

Terraform
Pulumi
OpenTelemetry
Prometheus/Grafana

Job description

uRun is building a universal runtime for real-time, stateful AI inference. We seek a senior engineer to design and own the scalable, low-latency infrastructure powering live, multi-user AI workloads. You’ll drive latency, memory usage, and QoS as a hands-on technical leader reporting to the founder.

You’ll shape platform standards, mentor engineers, and partner with ML and video teams to push the boundaries of real-time AI infrastructure in a fast-paced startup.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Founding Real-Time AI Platform Engineer
Founding Real-Time AI Platform Engineer

uRun • San Francisco (CA)

On-site
USD 140,000 - 180,000
Competitive salary and equity
Full health, dental, and vision coverage
401(k) retirement savings
+4
Founding Backend Engineer: Real-Time AI Runtime
Founding Backend Engineer: Real-Time AI Runtime

uRun • San Francisco (CA)

On-site
USD 120,000 - 160,000
Competitive salary and equity
Full health coverage
401(k) with company support
+3
Founding Engineer - Software
Founding Engineer - Software

uRun • San Francisco (CA)

On-site
USD 120,000 - 160,000
Competitive salary and equity
Full health coverage
401(k) with company support
+3
Founding ML Inference Performance Engineer
Founding ML Inference Performance Engineer

uRun • San Francisco (CA)

On-site
USD 120,000 - 160,000
Health, dental, and vision
401(k) participation
Flexible spending accounts
+3
Founding ML Infra Architect
Founding ML Infra Architect

uRun • San Francisco (CA)

On-site
USD 120,000 - 160,000
Health, dental, and vision
401(k)
Paid time off
+2
Founding Engineer - Platform
Founding Engineer - Platform

uRun • United States

Remote
USD 180,000 - 260,000
Equity
Health, dental, and vision
401(k)
+4
Founding Engineer, Real-Time AI Security Infra
Founding Engineer, Real-Time AI Security Infra

Recruiting From Scratch • San Francisco (CA)

On-site
USD 200,000 - 400,000
Significant equity
Visa transfer and sponsorship support
Founding Engineer, Real-Time AI Systems
Founding Engineer, Real-Time AI Systems

Cymon AI • United States

Hybrid
USD 120,000 - 160,000
Runtime Engineer: High-Performance AI Compute
Runtime Engineer: High-Performance AI Compute

SambaNova • Palo Alto (CA)

On-site
USD 110,000 - 150,000
95% premium coverage for employee medical insurance
Dental and Vision insurance
Health Savings Account with employer contribution
+2
Engine Architect for Long-Running AI Agent Runtime
Engine Architect for Long-Running AI Agent Runtime

Runta • San Mateo (CA)

On-site
USD 180,000 - 260,000