Senior ML Engineer - Real-Time Inference & Systems

Inworld AI

Germany (OH)

On-site

USD 120,000 - 180,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

A leading AI research lab is seeking innovative engineers to optimize and scale their realtime models. Candidates should possess a strong background in high-performance systems using C++, CUDA, and Kubernetes. Responsibilities include clarifying complex problems, contributing to significant systems programming efforts, and ensuring reliability in production. The position offers opportunities for relocation support to the San Francisco Bay Area.

Qualifications

  • Understanding of modern serving frameworks and techniques like vLLM or TRT-LLM.
  • Hands-on experience with quantization, distillation, and caching strategies.
  • Proficiency in coding optimization for performance on NVIDIA GPUs.
  • Experience with multi-GPU/multi-node inference.

Responsibilities

  • Make unclear problems clear and tackle performance challenges head-on.
  • Engage with US-based teams and contribute to non-trivial systems programming projects.
  • Ensure models are containerized, optimized for serving, and run reliably in production.

Skills

Inference Optimization
Model Acceleration
High-Performance Systems
Distributed Systems & Scaling
Public work
Full-cycle ownership
Professional fluency in English

Education

PhD in CS, Physics, Math, or equivalent practical experience

Tools

C++
CUDA
Rust
Python
Kubernetes
Ray

Job description

A leading AI research lab is seeking innovative engineers to optimize and scale their realtime models. Candidates should possess a strong background in high-performance systems using C++, CUDA, and Kubernetes. Responsibilities include clarifying complex problems, contributing to significant systems programming efforts, and ensuring reliability in production. The position offers opportunities for relocation support to the San Francisco Bay Area.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior ML Engineer - Real-Time Inference & Scalable Systems
Senior ML Engineer - Real-Time Inference & Scalable Systems

careers.bitkraft.vc - Jobboard • Germany (OH)

On-site
USD 120,000 - 160,000
Senior ML Engineer - Low-Latency Inference & Systems
Senior ML Engineer - Low-Latency Inference & Systems

Inworld • Germany (OH)

Hybrid
USD 120,000 - 160,000
Senior ML Infra Engineer — Real-Time, Low-Latency
Senior ML Infra Engineer — Real-Time, Low-Latency

LMArena • San Francisco (CA)

On-site
USD 170,000 - 260,000
Competitive compensation
Comprehensive health benefits
Opportunity to work on cutting-edge AI
ML Performance Engineer – Real-Time Inference
ML Performance Engineer – Real-Time Inference

Odyssey • Palo Alto (CA)

On-site
USD 130,000 - 160,000
Senior ML Infra Engineer - Real-Time AI Backends
Senior ML Infra Engineer - Real-Time AI Backends

Acceler8 Talent • San Francisco (CA)

Hybrid
USD 233,000 - 275,000
Senior ML Systems Engineer - Model Inference & Efficiency
Senior ML Systems Engineer - Model Inference & Efficiency

Cohere • New York (NY)

Hybrid
USD 100,000 - 150,000
Inclusive culture and work environment
Weekly lunch stipend, in-office lunches & snacks
Full health and dental benefits
+4
Staff ML Engineer: Build Ultra-Fast AI at Scale (Relocation)
Staff ML Engineer: Build Ultra-Fast AI at Scale (Relocation)

Inworld AI • Mountain View (CA)

On-site
USD 270,000 - 500,000
Staff AI Systems Engineer — Inference & RL
Staff AI Systems Engineer — Inference & RL

Together • San Francisco (CA)

On-site
USD 200,000 - 280,000
Health insurance
Startup equity
Competitive benefits
Realtime ML Inference Engineer — Scalable Serving
Realtime ML Inference Engineer — Scalable Serving

Yobi AI • New York (NY)

Remote
USD 120,000 - 150,000
Senior ML Systems Engineer — Inference & Scale
Senior ML Systems Engineer — Inference & Scale

Acceler8 Talent • San Francisco (CA), Northern (KY)

Hybrid
USD 170,000 - 240,000