LILT is building a live translation product, and this role focuses on the real-time speech translation backend. You will own the end-to-end path from live audio input to translated output, combining streaming speech recognition and adaptive machine translation in a low-latency system deployed on GPU Kubernetes infrastructure.
What you’ll do
- Build and run services that support high-throughput, real-time audio and text streaming.
- Own signal processing, session lifecycle management, and concurrency to keep the system stable under load.
- Integrate and serve streaming speech recognition and machine translation models, working with research teams to meet latency budgets.
- Develop model-driven confidence scoring and routing logic for human intervention when needed.
- Broadcast real-time updates and corrections to end users.
- Architect and scale production ML infrastructure on GPU-accelerated Kubernetes, including Ray Serve deployments.
- Implement batching, load balancing, and autoscaling strategies to maintain both performance and cost efficiency.
- Set up instrumentation to measure real-time performance and establish production-grade observability.
- Identify bottlenecks, optimize end-to-end throughput, and reduce end-to-end latency to meet production standards.
- Define technical contracts and interfaces for audio ingestion and downstream service integrations.
- Collaborate with frontend and platform engineering teams to maintain robust integration points.
- Drive cross-team alignment through clear API interfaces and technical contracts across engineering and product teams.
Key requirements
- BS or MS in Computer Science (or related) or equivalent practical experience.
- 3+ years building production backend or ML serving systems in Python, with strong async skills (asyncio).
- Hands-on experience with real-time streaming transport such as WebSocket or gRPC bidirectional streaming, including session state, backpressure, and connection lifecycle handling.
- Experience serving ML models on GPUs in production using Ray Serve, Triton, vLLM, or similar, along with Docker and Kubernetes.
- Experience integrating speech or NLP models into production systems, ideally streaming ASR with partial hypotheses, endpointing, and VAD.
- A latency-engineering mindset: you have profiled, instrumented, and optimized real-time or low-latency systems and can reason using per-stage budgets.
- Ability to use AI coding agents (Claude Code, Codex, or similar) with strong fundamentals: you can debug, review, and reason about every line and know when not to trust generated output.
- US citizenship and residence in the United States (contract requirement).
Technologies
- Python, asyncio, WebSocket, gRPC
- Ray Serve, Triton, vLLM
- Docker, Kubernetes
- Claude Code, Codex
- COMET, CometKiwi
- RabbitMQ
- Datadog, Prometheus
- WebRTC, SFU
- LiveKit Agents, Pipecat
Location and eligibility
- Boston, MA (onsite)
- Requires US citizenship and residence in the United States.
- Preferred locations include Washington, D.C.; Boston, MA; and Indianapolis, IN (East Coast / ET timezone preferred).
Compensation
USD 120,000 - 161,434 per yearly.
Preferred qualifications
- Ray Serve experience, including streaming responses and model multiplexing.
- Familiarity with simultaneous or incremental MT concepts (retranslation, prefix stability, wait-k policies).
- Machine translation quality estimation in the COMET/CometKiwi class, or other production confidence estimation.
- Message brokers for real-time fan-out and state distribution (RabbitMQ or similar).
- Streaming text-to-speech integration and time-to-first-audio optimization.
- WebRTC and SFU concepts, or voice pipeline frameworks such as LiveKit Agents or Pipecat.
- Handling CJK and other non-Latin text in NLP pipelines (Japanese, Korean, and English are first languages).
- Observability tooling for production ML systems (Datadog, Prometheus).