Staff ML Engineer: Build Ultra-Fast AI at Scale (Relocation)

Inworld AI

Mountain View (CA)

On-site

USD 270,000 - 500,000

Full time

14 days+
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Benefits offered by this job

Relocation assistance
Equity options
Comprehensive benefits package

Job summary

A tech-driven AI firm in Mountain View is seeking a high-performance systems engineer. This role demands expertise in inference optimization, model acceleration, and a deep understanding of performance tuning in C++, CUDA, and more. Candidates should possess a PhD or equivalent experience. Compensation ranges from $270,000 to $500,000 plus bonuses and equity. Relocation assistance is provided for successful candidates looking to innovate in a collaborative atmosphere.

Qualifications

  • Deep understanding of modern serving frameworks like vLLM or TRT-LLM.
  • Hands-on experience with quantization and caching strategies.
  • Proficiency in profiling code and optimizing performance for NVIDIA GPUs.

Responsibilities

  • Take models from research, containerize and optimize for production.
  • Communicate and collaborate closely with the team.
  • Design prototypes to explore unclear problems.

Skills

Inference Optimization
Model Acceleration
High-Performance Systems
Distributed Systems & Scaling
Public work
Full-cycle ownership
PhD in CS, Physics, Math or equivalent

Education

PhD in Computer Science or equivalent

Tools

C++
CUDA
Rust
Python
Kubernetes

Job description

A tech-driven AI firm in Mountain View is seeking a high-performance systems engineer. This role demands expertise in inference optimization, model acceleration, and a deep understanding of performance tuning in C++, CUDA, and more. Candidates should possess a PhD or equivalent experience. Compensation ranges from $270,000 to $500,000 plus bonuses and equity. Relocation assistance is provided for successful candidates looking to innovate in a collaborative atmosphere.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Staff ML Engineer — Ultra-Low-Latency Inference
Staff ML Engineer — Ultra-Low-Latency Inference

Inworld • Mountain View (CA)

Hybrid
USD 270,000 - 500,000
Relocation assistance
Equity options
Comprehensive benefits
Senior ML Systems Engineer - Model Inference & Efficiency
Senior ML Systems Engineer - Model Inference & Efficiency

Cohere • New York (NY)

Hybrid
USD 100,000 - 150,000
Inclusive culture and work environment
Weekly lunch stipend, in-office lunches & snacks
Full health and dental benefits
+4
Senior ML Services Engineer — Onsite in Mountain View
Senior ML Services Engineer — Onsite in Mountain View

AI Fund • Mountain View (CA)

On-site
USD 190,000 - 220,000
Software Engineer - ML Model Performance
Software Engineer - ML Model Performance

Baseten • San Francisco (CA)

On-site
USD 150,000 - 250,000
Staff ML Performance Engineer: Scale Training Throughput
Staff ML Performance Engineer: Scale Training Throughput

Wayve • Sunnyvale (CA)

On-site
USD 130,000 - 160,000
Engineering Manager, ML Inference & Scale
Engineering Manager, ML Inference & Scale

Anthropic • San Francisco (CA)

Hybrid
USD 425,000 - 560,000
Competitive compensation
Flexible working hours
Generous vacation and parental leave
Inference Optimization Engineer: Fast, Cost-Effective ML
Inference Optimization Engineer: Fast, Cost-Effective ML

Build AI • San Francisco (CA)

On-site
USD 150,000 - 210,000
Competitive pay
Medical, dental, and vision packages
Housing subsidy $2k/month near SF offi
+6
Senior AI Systems Performance Engineer: Drive SOTA Inference
Senior AI Systems Performance Engineer: Drive SOTA Inference

SambaNova • Palo Alto (CA)

On-site
USD 120,000 - 150,000
95% premium coverage for employee medical insurance
Health Savings Account with employer contribution
Flexible Spending Account options
AI/ML Engineer: Next‑Gen Platforms & GPU Workloads
AI/ML Engineer: Next‑Gen Platforms & GPU Workloads

VeeAR Projects Inc. • Sunnyvale (CA)

On-site
USD 140,000 - 210,000
ML Performance Engineer – Real-Time Inference
ML Performance Engineer – Real-Time Inference

Odyssey • Palo Alto (CA)

On-site
USD 130,000 - 160,000