Machine Learning Engineer (Real-Time Speech Translation)

Lilt

United States

On-site

USD 140,000 - 190,000

Full time

4 days ago
Be an early applicant
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Benefits offered by this job

Medical benefits
401(k) matching
Equity

Job summary

Lilt is hiring an ML Engineer to own the real-time speech translation backend end-to-end. You will connect live audio input to translated output, building on Ray Serve on GPU Kubernetes clusters and in-house MT models.

You’ll work with architects and researchers to achieve low latency and reliable production service. The role requires Python expertise, real-time streaming experience, and the ability to optimize ML stack from models to infrastructure, with a strong emphasis on cost‑aware hardware

Qualifications

  • 3+ years building production backend or ML serving systems in Python
  • Experience integrating speech or NLP models into production systems, ideally streaming ASR
  • US citizenship and residence in the United States (contract requirement)
  • Hands‑on experience with real‑time streaming transport: WebSocket or gRPC bidirectional streaming
  • A latency‑engineering mindset: you have profiled, instrumented, and optimized a real‑time or low‑latency system
  • Experience serving ML models in production on GPUs (Ray Serve, Triton, vLLM) with Docker and Kubernetes
  • Effective use of AI coding agents on top of fundamentals learned the hard way
  • Observability tooling (Datadog, Prometheus) for production ML systems
  • Handling of CJK and other non-Latin text in NLP pipelines
  • Streaming text-to-speech integration and time-to-first-audio optimization
  • Machine translation quality estimation or confidence estimation in production
  • Familiarity with streaming data architectures and message brokers

Responsibilities

  • Build the real-time speech translation backend end-to-end, from live audio input to translated output
  • Integrate and serve streaming ASR and MT models within latency budgets
  • Architect and scale production ML infrastructure on GPU‑accelerated Kubernetes clusters
  • Implement batching, load balancing, and autoscaling for performance and cost‑efficiency
  • Define API contracts for audio ingestion and downstream services
  • Collaborate with product and research teams to meet latency targets
  • Drive cross‑team alignment and code reviews using AI agents responsibly

Skills

Python proficiency
Async programming
Latency engineering
Production backend
Real-time streaming
Observability tooling
AI coding agents

Education

Bachelor's or Master’s in CS or related

Tools

WebSocket
gRPC
Docker
Kubernetes
Ray Serve
Triton
vLLM
LiveKit
RabbitMQ

Job description

  • We are building a new live translation product
  • We are looking for an ML Engineer to build the real-time speech translation backend that powers it
  • You will own the real-time speech translation backend end-to-end, from live audio input to translated output
  • You will build on LILT’s production model serving platform (Ray Serve on GPU Kubernetes clusters) and our in‑house adaptive machine translation models, working closely with the senior architects of that platform and with our language processing researchers
  • The ASR and MT models exist
  • Your job is to make them work together as a low‑latency streaming system that holds up in production
  • This is a hands‑on backend engineering role for those looking to own a real‑time ML system from end to end, supported by expert guidance and a clear product vision
  • Our team adopts an AI‑first approach, leveraging agentic coding and AI‑driven PR reviews to accelerate development
  • We combine this with deep technical expertise, requiring not only expert Python proficiency but also a comprehensive understanding of the entire ML stack, from optimizing neural network architectures to managing production infrastructure on Kubernetes, and making informed, cost‑aware decisions on hardware selection
  • Real‑time pipeline architecture: Build and manage services for high‑throughput, real‑time audio and text streaming. Handle signal processing, session lifecycles, and concurrency management to ensure robust operation under load
  • ML model integration: Integrate and serve streaming speech recognition and machine translation models, collaborating with research teams to ensure models operate within required latency budgets
  • Quality and confidence workflows: Develop logic for model‑based confidence scoring, routing segments for human intervention as needed, and broadcasting real‑time updates and corrections to end‑users
  • Infrastructure and scale: Architect and scale production ML infrastructure on GPU‑accelerated Kubernetes clusters. Implement batching, load balancing, and autoscaling strategies to maintain performance and cost‑efficiency
  • Latency engineering: Establish comprehensive instrumentation for real‑time performance. Identify bottlenecks, optimize system throughput, and drive down end‑to‑end latency metrics to meet production standards
  • Interface and API definition: Define technical contracts and interfaces for audio ingestion and downstream service integrations. Partner with frontend and platform engineering teams to maintain clean, robust integration points
  • Collaboration and technical leadership: Drive cross‑team alignment by defining clear API interfaces and technical contracts, facilitating effective communication between engineering and product teams to ensure seamless system integration
Benefits
  • Medical Benefits: Employees receive coverage of medical, dental, and vision insurance, plus FSA/DFSA, HSA, and Commuter benefits. In addition, LILT pays for basic life insurance, short‑term disability, and long‑term disability
  • Paid parental leave is provided after 6 months
  • Monthly lifestyle benefit stipend via the Fringe platform to allow employees to customize benefits to their lifestyle
  • Compensation: Meaningful equity, 401(k) matching, and flexible time off plus company holidays

3+ years building production backend or ML serving systems in Python, including strong async programming (asyncio) skillsExperience integrating speech or NLP models into production systems, ideally streaming ASR (partial hypotheses, endpointing, VAD)US citizenship and residence in the United States (contract requirement)BS or MS in Computer Science or a related field, or equivalent practical experienceHands‑on experience with real‑time streaming transport: WebSocket or gRPC bidirectional streaming, session state, backpressure, and connection lifecycle handlingA latency‑engineering mindset: you have profiled, instrumented, and optimized a real‑time or low‑latency system and can reason in per‑stage budgetsExperience serving ML models in production on GPUs (Ray Serve, Triton, vLLM, or similar), with Docker and KubernetesEffective use of AI coding agents (Claude Code, Codex, or similar) on top of fundamentals learned the hard way: you let agents do the typing, but you can debug, review, and reason about every line without them, and you know when not to trust themObservability tooling (Datadog, Prometheus) for production ML systemsHandling of CJK and other non‑Latin text in NLP pipelines (our first languages are Japanese, Korean, and English)WebRTC and SFU concepts, or voice pipeline frameworks (LiveKit Agents, Pipecat)Streaming text‑to‑speech integration and time‑to‑first‑audio optimizationMachine translation quality estimation (COMET/CometKiwi class models) or other confidence estimation in productionMessage brokers for real‑time fan‑out and state distribution (RabbitMQ or similar)Familiarity with simultaneous or incremental MT concepts (retranslation, prefix stability, wait‑k policies)Ray Serve specifically, including streaming responses and model multiplexing

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Machine Learning Engineer (Real-Time Speech Translation)
Machine Learning Engineer (Real-Time Speech Translation)

LILT • Boston (MA)

On-site
USD 120,000 - 161,000
Machine Learning Engineer (Real-Time Speech Translation)
Machine Learning Engineer (Real-Time Speech Translation)

LILT • Washington

On-site
USD 140,000 - 200,000
Machine Learning Engineer (Real-Time Speech Translation)
Machine Learning Engineer (Real-Time Speech Translation)

lilt-corporate • Washington (IN)

On-site
USD 140,000 - 210,000
Real-Time Speech Translation ML Engineer
Real-Time Speech Translation ML Engineer

lilt-corporate • Washington (IN)

On-site
USD 140,000 - 210,000
Real-Time Speech Translation Engineer
Real-Time Speech Translation Engineer

LILT • Washington

On-site
USD 140,000 - 200,000
Real-Time Speech Translation Backend Engineer
Real-Time Speech Translation Backend Engineer

Lilt • United States

On-site
USD 140,000 - 190,000
Medical benefits
401(k) matching
Equity
Staff Fullstack Engineer - Internal Tools
Staff Fullstack Engineer - Internal Tools

lilt-corporate • United States

On-site
USD 120,000 - 180,000
ML Engineer — Real-time Speech
ML Engineer — Real-time Speech

Sellsig • Minnesota

On-site
USD 100,000 - 130,000
Senior Full Stack Engineer
Senior Full Stack Engineer

Lilt • United States

On-site
USD 180,000 - 240,000
Medical benefits
Parental leave
Fringe stipend
+3
Real-Time Speech Translation Backend Engineer
Real-Time Speech Translation Backend Engineer

LILT • Boston (MA)

On-site
USD 120,000 - 161,000