Senior ML Engineer - Real-Time Inference & Systems
Inworld AI
Germany (OH)
On-site
USD 120,000 - 180,000
Full time
14 days+
Get more replies from employers
Send a job-specific resume in minutes.
Start fresh or import an existing resume
Job summary
A leading AI research lab is seeking innovative engineers to optimize and scale their realtime models. Candidates should possess a strong background in high-performance systems using C++, CUDA, and Kubernetes. Responsibilities include clarifying complex problems, contributing to significant systems programming efforts, and ensuring reliability in production. The position offers opportunities for relocation support to the San Francisco Bay Area.
Qualifications
Understanding of modern serving frameworks and techniques like vLLM or TRT-LLM.
Hands-on experience with quantization, distillation, and caching strategies.
Proficiency in coding optimization for performance on NVIDIA GPUs.
Experience with multi-GPU/multi-node inference.
Responsibilities
Make unclear problems clear and tackle performance challenges head-on.
Engage with US-based teams and contribute to non-trivial systems programming projects.
Ensure models are containerized, optimized for serving, and run reliably in production.
Skills
Inference Optimization
Model Acceleration
High-Performance Systems
Distributed Systems & Scaling
Public work
Full-cycle ownership
Professional fluency in English
Education
PhD in CS, Physics, Math, or equivalent practical experience
Tools
C++
CUDA
Rust
Python
Kubernetes
Ray
Job description
A leading AI research lab is seeking innovative engineers to optimize and scale their realtime models. Candidates should possess a strong background in high-performance systems using C++, CUDA, and Kubernetes. Responsibilities include clarifying complex problems, contributing to significant systems programming efforts, and ensuring reliability in production. The position offers opportunities for relocation support to the San Francisco Bay Area.