Senior ML Engineer - Low-Latency Inference & Systems

Inworld

Germany (OH)

Hybrid

USD 120,000 - 160,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

A leading AI technology firm in the United States is seeking an experienced engineer to optimize model performance. The role requires expertise in inference optimization, model acceleration, and proficiency in C++, CUDA, and Python, among other skills. You'll work collaboratively with teams to tackle complex problems, ensuring high-quality outcomes. Professional fluency in English is essential for daily collaboration. The company provides potential relocation support for candidates interested in the San Francisco Bay Area.

Qualifications

  • Strong understanding of modern serving frameworks and techniques.
  • Hands-on experience with quantization, distillation, and continuous batching.
  • Proficiency in C++, CUDA, Rust, or optimised Python.

Responsibilities

  • Optimize model serving and ensure reliability in production.
  • Collaborate with US-based teams on unclear problems.
  • Take ownership from research to production.

Skills

Inference Optimization
Model Acceleration
High-Performance Systems
Distributed Systems & Scaling
Public work
Full-cycle ownership
Background in CS, Physics, or Math
Professional fluency in English

Education

PhD in CS, Physics, Math, or equivalent

Job description

A leading AI technology firm in the United States is seeking an experienced engineer to optimize model performance. The role requires expertise in inference optimization, model acceleration, and proficiency in C++, CUDA, and Python, among other skills. You'll work collaboratively with teams to tackle complex problems, ensuring high-quality outcomes. Professional fluency in English is essential for daily collaboration. The company provides potential relocation support for candidates interested in the San Francisco Bay Area.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior ML Engineer - Real-Time Inference & Systems
Senior ML Engineer - Real-Time Inference & Systems

Inworld AI • Germany (OH)

On-site
USD 120,000 - 180,000
Senior ML Systems Engineer - Model Inference & Efficiency
Senior ML Systems Engineer - Model Inference & Efficiency

Cohere • New York (NY)

Hybrid
USD 100,000 - 150,000
Inclusive culture and work environment
Weekly lunch stipend, in-office lunches & snacks
Full health and dental benefits
+4
Staff ML Engineer: Efficient ML & Low-Latency AI
Staff ML Engineer: Efficient ML & Low-Latency AI

Embedding VC • San Francisco (CA)

On-site
USD 100,000 - 150,000
Software Engineer - ML Model Performance
Software Engineer - ML Model Performance

Baseten • San Francisco (CA)

On-site
USD 150,000 - 250,000
Senior ML Engineer - Real-Time Inference & Scalable Systems
Senior ML Engineer - Real-Time Inference & Scalable Systems

careers.bitkraft.vc - Jobboard • Germany (OH)

On-site
USD 120,000 - 160,000
Senior AI Model Serving Engineer — Low-Latency Inference
Senior AI Model Serving Engineer — Low-Latency Inference

Menlo Ventures • San Francisco (CA)

On-site
USD 166,000 - 225,000
Annual performance bonus
Equity options
Comprehensive benefits package
Staff ML Engineer: Build Ultra-Fast AI at Scale (Relocation)
Staff ML Engineer: Build Ultra-Fast AI at Scale (Relocation)

Inworld AI • Mountain View (CA)

On-site
USD 270,000 - 500,000
Staff Engineer - ML Inference & Model Efficiency
Staff Engineer - ML Inference & Model Efficiency

Cohere • San Francisco (CA)

Remote
USD 180,000 - 240,000
Inclusive work culture
Weekly lunch stipend
Full health and dental benefits
+4
Staff ML Infrastructure & Performance Engineer
Staff ML Infrastructure & Performance Engineer

Embedding VC • San Mateo (CA)

On-site
USD 120,000 - 150,000
Senior Inference Performance Engineer — GPU & CUDA
Senior Inference Performance Engineer — GPU & CUDA

Inference • San Francisco (CA)

Hybrid
USD 220,000 - 320,000
Competitive compensation
Equity in a high-growth startup
Comprehensive benefits