Real-Time Speech Translation Backend Engineer

Lilt

United States

On-site

USD 140,000 - 190,000

Full time

3 days ago
Be an early applicant
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Benefits offered by this job

Medical benefits
401(k) matching
Equity

Job summary

Lilt is hiring an ML Engineer to own the real-time speech translation backend end-to-end. You will connect live audio input to translated output, building on Ray Serve on GPU Kubernetes clusters and in-house MT models.

You’ll work with architects and researchers to achieve low latency and reliable production service. The role requires Python expertise, real-time streaming experience, and the ability to optimize ML stack from models to infrastructure, with a strong emphasis on cost‑aware hardware

Qualifications

  • 3+ years building production backend or ML serving systems in Python
  • Experience integrating speech or NLP models into production systems, ideally streaming ASR
  • US citizenship and residence in the United States (contract requirement)
  • Hands‑on experience with real‑time streaming transport: WebSocket or gRPC bidirectional streaming
  • A latency‑engineering mindset: you have profiled, instrumented, and optimized a real‑time or low‑latency system
  • Experience serving ML models in production on GPUs (Ray Serve, Triton, vLLM) with Docker and Kubernetes
  • Effective use of AI coding agents on top of fundamentals learned the hard way
  • Observability tooling (Datadog, Prometheus) for production ML systems
  • Handling of CJK and other non-Latin text in NLP pipelines
  • Streaming text-to-speech integration and time-to-first-audio optimization
  • Machine translation quality estimation or confidence estimation in production
  • Familiarity with streaming data architectures and message brokers

Responsibilities

  • Build the real-time speech translation backend end-to-end, from live audio input to translated output
  • Integrate and serve streaming ASR and MT models within latency budgets
  • Architect and scale production ML infrastructure on GPU‑accelerated Kubernetes clusters
  • Implement batching, load balancing, and autoscaling for performance and cost‑efficiency
  • Define API contracts for audio ingestion and downstream services
  • Collaborate with product and research teams to meet latency targets
  • Drive cross‑team alignment and code reviews using AI agents responsibly

Skills

Python proficiency
Async programming
Latency engineering
Production backend
Real-time streaming
Observability tooling
AI coding agents

Education

Bachelor's or Master’s in CS or related

Tools

WebSocket
gRPC
Docker
Kubernetes
Ray Serve
Triton
vLLM
LiveKit
RabbitMQ

Job description

Lilt is hiring an ML Engineer to own the real-time speech translation backend end-to-end. You will connect live audio input to translated output, building on Ray Serve on GPU Kubernetes clusters and in-house MT models.

You’ll work with architects and researchers to achieve low latency and reliable production service. The role requires Python expertise, real-time streaming experience, and the ability to optimize ML stack from models to infrastructure, with a strong emphasis on cost‑aware hardware

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Real-Time Speech Translation ML Engineer
Real-Time Speech Translation ML Engineer

lilt-corporate • Washington (IN)

On-site
USD 140,000 - 210,000
Real-Time Speech Translation Engineer
Real-Time Speech Translation Engineer

LILT • Washington

On-site
USD 140,000 - 200,000
Real-Time Speech Translation Backend Engineer
Real-Time Speech Translation Backend Engineer

LILT • Boston (MA)

On-site
USD 120,000 - 161,000
Machine Learning Engineer (Real-Time Speech Translation)
Machine Learning Engineer (Real-Time Speech Translation)

Lilt • United States

On-site
USD 140,000 - 190,000
Medical benefits
401(k) matching
Equity
Real-Time Translation Backend Engineer
Real-Time Translation Backend Engineer

Lilt • United States

On-site
USD 120,000 - 180,000
Medical, dental, vision insurance
Paid parental leave
Lifestyle stipend via Fringe platform
+1
Machine Learning Engineer (Real-Time Speech Translation)
Machine Learning Engineer (Real-Time Speech Translation)

lilt-corporate • Washington (IN)

On-site
USD 140,000 - 210,000
Machine Learning Engineer (Real-Time Speech Translation)
Machine Learning Engineer (Real-Time Speech Translation)

LILT • Washington

On-site
USD 140,000 - 200,000
Machine Learning Engineer (Real-Time Speech Translation)
Machine Learning Engineer (Real-Time Speech Translation)

LILT • Boston (MA)

On-site
USD 120,000 - 161,000
Senior Full-Stack Engineer — Real-Time AI Translation
Senior Full-Stack Engineer — Real-Time AI Translation

LILT • Washington

On-site
USD 180,000 - 240,000
Real-Time Translation Backend Engineer
Real-Time Translation Backend Engineer

lilt-corporate • Washington (IN)

On-site
USD 120,000 - 180,000