Get more replies from employers
Send a job-specific resume in minutes.
uRun in San Francisco is seeking a backend software engineer to build services for their real-time inference platform. This role focuses on low-latency systems and system architecture, working closely with product and applied AI teams.
The ideal candidate should have over 7 years of experience in backend development, with proficiency in languages like TypeScript and Python. Competitive salary, full health benefits, and a flexible work environment are offered.
AI inference today is slow, expensive, and stateless. Send a query, wait, get a response, reset. That's fine for batch, but AI is becoming interactive, and interactive means inference has to respond instantly, hold context across a session, and be steerable in real time.
Nobody had built infrastructure that does all three at once. The bottleneck isn't the models. It's the runtime underneath them.
uRun — Universal Runtime is the layer that makes real-time, stateful inference possible. Our platform lets AI respond instantly, hold context across a session, and be directed as it runs.
We prove it through the hardest problem in the stack: real-time AI video generation. Not pre-rendered clips. Not queued jobs. Live, steerable, continuous video that responds as you speak. Solve that, and the rest of the inference stack follows — and that's what we've done. We're an infrastructure company; we build the layer model labs, builders, and research teams ship on top of.
You'll build the services, APIs, and core application systems that power uRun's runtime the software layer that turns our real-time inference platform into something product and applied AI teams can actually build on.
This is not a conventional CRUD backend role. The work centres on low-latency, high-throughput systems: real-time interaction, evolving session state, and request handling that stays reliable under heavy compute and concurrency. You'll work closely with product, infrastructure, and applied AI teams, in an early-stage environment where the architecture is still being set.
You'd join an early team building a genuinely new category of real‑time AI infrastructure and product. You'd work directly with the founder and senior technical leadership, with real influence over architecture, developer velocity, and the shape of the product itself, working at the intersection of backend systems, platform engineering, and interactive AI.
It also means ambiguity. We're early: there's limited scaffolding, priorities shift, and you'll often be setting architecture before there's a playbook to follow. That's a real part of the job, and it suits some people far more than others. We'd rather you know it going in.