Software Engineer, ML Serving

Unusual Ventures

San Francisco (CA)

On-site

USD 180,000 - 240,000

Full time

6 days ago
Be an early applicant
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Job summary

Rime is hiring a Software Engineer to own the serving infrastructure that connects inference engines to production. You will work at the intersection of ML systems and cloud infrastructure, building, hardening, and scaling real-time voice model serving.

You'll own the TTS serving stack, optimize multi-node serving, and ensure cross-hardware compatibility from HPC GPUs to cloud deployments. You will shape architecture, CI/CD, and reliability while collaborating with ML teams.

Qualifications

  • Hands-on experience with real-time multinode ML serving infrastructure.
  • Experience with distributed or disaggregated model serving.
  • Strong cloud infrastructure fundamentals and IaC tooling.

Responsibilities

  • Architect and implement the TTS serving infrastructure across GPUs and API surface.
  • Optimize models from single-node to disaggregated fleet serving.
  • Ensure compatibility across NVIDIA hardware for on-prem and cloud deployments.
  • Maintain CI/CD workflows for the model serving pipeline.
  • Manage site reliability: on-call, monitoring, and observability.
  • Plan and manage GPU resource provisioning and costs.

Skills

Real-time ML serving
NVIDIA Dynamo/Triton
vLLM
SGLang
Distributed model serving
Cloud infrastructure
Linux
Docker
Kubernetes
Terraform

Tools

Docker
Kubernetes
Terraform
Packer

Job description

Software Engineer, ML Serving - Rime Ai

Rime is a foundation modeling company that builds voice AI for enterprises running customer experiences at scale. Our models are purpose-built for high-volume conversational deployments, engineered for the accuracy, performance, and deployment flexibility that production environments actually demand.

We started from a different premise than the rest of the field: build voice AI for human connection, not slop. Before we trained a single model, we built our own corpus: full-duplex, studio-quality conversational speech of normal people, recorded and annotated by linguists. It's why our models are unparalleled in naturalism, and it's why enterprises pick Rime when pilots need to make it to production.

Role Overview

We're hiring a Software Engineer to own the serving infrastructure that connects Rime's inference engines to the world. This role sits at the intersection of ML systems and cloud infrastructure — you'll work directly on model inference and cloud infrastructure to build, harden, and scale the systems that stream voice at real-time latency. As Rime moves toward its next-generation architecture, you'll be a core architect of how our models get served.

What You'll Own
  • Architecture and implementation of Rime's TTS serving infrastructure, from GPU-backed inference engines to the API surface.
  • Model optimization from a single-node to disaggregated fleet serving.
  • Compatibility with different NVIDIA hardwares from Hopper to Blackwell and beyond for on-prem and cloud deployments.
  • Continuous integration and deployment workflows for the model serving pipeline.
  • Site reliability: on-call rotation, monitoring, alerting, and observability across the serving stack.
  • Resource provision, cost management across our GPU fleet.
What We're Looking For
  • Hands-on experience with real-time multinode ML serving infrastructure — ML serving framework experience: NVIDIA Dynamo/Triton, vLLM, SGLang, or equivalent.
  • Experience with distributed or disaggregated model serving (Tensor Parallel, Pipeline Parallel, or equivalent).
  • Strong cloud infrastructure fundamentals: Linux internals, networking, containerization (Docker, Kubernetes).
  • IaC experience — Terraform, Packer, or comparable. You should have opinions about how to do this right.
  • On-call is part of the job. You treat production reliability as a shared responsibility.
Nice to Have
  • Experience with multinode training (DDP, FSDP, etc.).
  • Experience with gRPC or other bidirectional binary streaming protocols.
  • Experience with audio streaming and related technologies (WebRTC, WebSockets, etc.).
  • Experience with a multilingual monorepo where you pick the best language out of merit more than personal experience.
  • Experience with multi-cloud infrastructures (AWS, GCP, OCI, etc.).
  • Comfort with configuration management tooling (Ansible, Chef, Puppet, or similar).
  • SRE, DevOps, or platform engineering background at a startup.
  • Experience at an early-stage company.
Why Join Rime
  • Build the serving infrastructure behind a category-defining voice AI company from the ground up.
  • You will bring in experience that no one else currently has at the company: you can help us set the vision.
  • Direct collaboration with the inference, platform, and ML teams — no handoff culture.
  • The systems you build determine what experiences our customers can deploy at scale.
  • Meaningful equity upside at an early stage.
  • High ownership, high standards, low bureaucracy.
  • SF / Bay Area.
At Rime, we...
  • Are outliers
  • Cut through the hype to focus on the craft
  • Move fast with agency and freedom
  • Maintain a growth mindset, finding joy in the struggle
  • Do the right things, knowing that it'll lead to making money

If that sounds like you too, you'll be a great fit for Rime!

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Machine Learning Scientist
Machine Learning Scientist

Rime • San Francisco (CA)

Hybrid
USD 180,000 - 240,000
Remote-friendly
Visa sponsorship available
Equity upside
+1
Forward Deployed Engineer
Forward Deployed Engineer

Rime Labs • San Francisco (CA)

Hybrid
USD 185,000 - 235,000
Competitive compensation
Equity options
Exposure to cutting-edge AI technology
Software Engineer, Growth - Rime Ai
Software Engineer, Growth - Rime Ai

Unusual Ventures • San Francisco (CA)

On-site
USD 120,000 - 160,000
Meaningful equity
High ownership
Close collaboration with founders
ML Serving Engineer – Real-Time Voice AI Platform
ML Serving Engineer – Real-Time Voice AI Platform

Unusual Ventures • San Francisco (CA)

On-site
USD 180,000 - 240,000
Founding Engineer - ML Infrastructure
Founding Engineer - ML Infrastructure

uRun • San Francisco (CA)

On-site
USD 120,000 - 160,000
Health, dental, and vision
401(k)
Paid time off
+2
ML Ops Infrastructure Engineer
ML Ops Infrastructure Engineer

Deepgram • United States

Remote
USD 140,000 - 180,000
Medical, dental, vision benefits
Unlimited PTO
401(k) plan with company match
+1
Staff Machine Learning Engineer, Voice AI
Staff Machine Learning Engineer, Voice AI

Together • San Francisco (CA)

On-site
USD 220,000 - 280,000
Member of Technical Staff — Model Optimization and Inference (New Grad)
Member of Technical Staff — Model Optimization and Inference (New Grad)

Nuance Labs • Seattle (WA)

On-site
USD 200,000 - 300,000
Health Savings Account with $2,000 annual contributions
15 days of PTO plus public holidays
Lunch, drinks, and snacks provided daily
Inference Infrastructure Engineer, Serving
Inference Infrastructure Engineer, Serving

Elorian • Palo Alto (CA)

On-site
USD 200,000 - 400,000
Health, dental, and vision benefits
Unlimited PTO
Parental leave
+1
Hands-On Forward-Deployed Voice AI Engineer
Hands-On Forward-Deployed Voice AI Engineer

Rime Labs • San Francisco (CA)

Hybrid
USD 185,000 - 235,000
Competitive compensation
Equity options
Exposure to cutting-edge AI technology