Senior ML Engineer - Real-Time Inference & Scalable Systems
careers.bitkraft.vc - Jobboard
Germany (OH)
On-site
USD 120,000 - 160,000
Full time
14 days+
Get more replies from employers
Send a job-specific resume in minutes.
Start fresh or import an existing resume
Job summary
Inworld is seeking an experienced ML Engineer in the United States. You will work on high-performance multimodal models and lead the optimization of performance in production systems. Candidates should have solid experience with C++, CUDA, Kubernetes, and a deep understanding of machine learning frameworks. A PhD in a related field is preferred. This role involves close collaboration with leadership teams and may include relocation support to the San Francisco Bay Area.
Qualifications
Deep understanding of modern serving frameworks and techniques like vLLM or TRT-LLM.
Hands-on experience with quantization, distillation, caching strategies, continuous batching.
Proficiency in C++, CUDA, Rust, or optimized Python.
Experience with Kubernetes and multi-GPU inference.
Responsibilities
Develop high-performance models under real-time constraints.
Collaborate with teams on implementing scalable ML solutions.
Optimize model performance and ensure reliability in production.
Skills
C++
CUDA
Rust
Python
Kubernetes
Deep Learning
Performance Optimization
ML Systems
Education
PhD in Computer Science, Physics or Math
Job description
Inworld is seeking an experienced ML Engineer in the United States. You will work on high-performance multimodal models and lead the optimization of performance in production systems. Candidates should have solid experience with C++, CUDA, Kubernetes, and a deep understanding of machine learning frameworks. A PhD in a related field is preferred. This role involves close collaboration with leadership teams and may include relocation support to the San Francisco Bay Area.