Senior ML Engineer - Real-Time Inference & Scalable Systems

careers.bitkraft.vc - Jobboard

Germany (OH)

On-site

USD 120,000 - 160,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Inworld is seeking an experienced ML Engineer in the United States. You will work on high-performance multimodal models and lead the optimization of performance in production systems. Candidates should have solid experience with C++, CUDA, Kubernetes, and a deep understanding of machine learning frameworks. A PhD in a related field is preferred. This role involves close collaboration with leadership teams and may include relocation support to the San Francisco Bay Area.

Qualifications

  • Deep understanding of modern serving frameworks and techniques like vLLM or TRT-LLM.
  • Hands-on experience with quantization, distillation, caching strategies, continuous batching.
  • Proficiency in C++, CUDA, Rust, or optimized Python.
  • Experience with Kubernetes and multi-GPU inference.

Responsibilities

  • Develop high-performance models under real-time constraints.
  • Collaborate with teams on implementing scalable ML solutions.
  • Optimize model performance and ensure reliability in production.

Skills

C++
CUDA
Rust
Python
Kubernetes
Deep Learning
Performance Optimization
ML Systems

Education

PhD in Computer Science, Physics or Math

Job description

Inworld is seeking an experienced ML Engineer in the United States. You will work on high-performance multimodal models and lead the optimization of performance in production systems. Candidates should have solid experience with C++, CUDA, Kubernetes, and a deep understanding of machine learning frameworks. A PhD in a related field is preferred. This role involves close collaboration with leadership teams and may include relocation support to the San Francisco Bay Area.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior ML Engineer - Real-Time Inference & Systems
Senior ML Engineer - Real-Time Inference & Systems

Inworld AI • Germany (OH)

On-site
USD 120,000 - 180,000
Senior ML Engineer - Low-Latency Inference & Systems
Senior ML Engineer - Low-Latency Inference & Systems

Inworld • Germany (OH)

Hybrid
USD 120,000 - 160,000
Staff ML Performance Engineer — Scalable Inference & CUDA
Staff ML Performance Engineer — Scalable Inference & CUDA

Modal • New York (NY)

On-site
USD 120,000 - 160,000
Senior ML Infra Engineer — Real-Time, Low-Latency
Senior ML Infra Engineer — Real-Time, Low-Latency

LMArena • San Francisco (CA)

On-site
USD 170,000 - 260,000
Competitive compensation
Comprehensive health benefits
Opportunity to work on cutting-edge AI
Senior ML Engineer, Voice AI — Real-Time Inference Lead
Senior ML Engineer, Voice AI — Real-Time Inference Lead

Together AI • San Francisco (CA)

On-site
USD 200,000 - 260,000
ML Performance Engineer – Real-Time Inference
ML Performance Engineer – Real-Time Inference

Odyssey • Palo Alto (CA)

On-site
USD 130,000 - 160,000
Senior ML Systems Engineer - Model Inference & Efficiency
Senior ML Systems Engineer - Model Inference & Efficiency

Cohere • New York (NY)

Hybrid
USD 100,000 - 150,000
Inclusive culture and work environment
Weekly lunch stipend, in-office lunches & snacks
Full health and dental benefits
+4
Senior ML Platform Engineer - Remote, Scalable Inference
Senior ML Platform Engineer - Remote, Scalable Inference

PlayStation Global • San Diego (CA)

Hybrid
USD 120,000 - 160,000
Realtime ML Inference Engineer — Scalable Serving
Realtime ML Inference Engineer — Scalable Serving

Yobi AI • New York (NY)

Remote
USD 120,000 - 150,000
Senior / Lead Machine Learning Engineer, Serving - Germany
Senior / Lead Machine Learning Engineer, Serving - Germany

careers.bitkraft.vc - Jobboard • Germany (OH)

On-site
USD 120,000 - 160,000