Staff ML Engineer — Ultra-Low-Latency Inference

Inworld

Mountain View (CA)

Hybrid

USD 270,000 - 500,000

Full time

14 days+
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Benefits offered by this job

Relocation assistance
Equity options
Comprehensive benefits

Job summary

A tech company in Mountain View seeks talented engineers for a role emphasizing high-performance systems, inference optimization, and model acceleration. You will thrive in ambiguity, tackle unclear problems, and design impactful solutions. The position offers a competitive base salary between $270,000 and $500,000, along with bonuses, equity, and benefits. Ideal candidates hold a relevant PhD or equivalent experience and have a knack for full-cycle ownership in engineering projects. Join a dynamic team prioritizing performance and collaboration.

Qualifications

  • Deep understanding of inference optimization frameworks.
  • Hands-on experience with model acceleration techniques.
  • Proficiency in C++, CUDA, Rust, or optimized Python.
  • Experience with Kubernetes and multi-GPU inference.
  • Track record of impactful public work in systems programming.
  • Ability to own models from research to production.

Responsibilities

  • Transform unclear problems into clear solutions.
  • Prioritize impact over theoretical optimizations.
  • Ensure stability and performance before launch.
  • Foster collaboration and strong team culture.

Skills

Inference Optimization
Model Acceleration
High-Performance Systems
Distributed Systems & Scaling
Public work
Full-cycle ownership
PhD in CS/Physics/Math

Education

PhD in CS, Physics, Math, or equivalent experience

Job description

A tech company in Mountain View seeks talented engineers for a role emphasizing high-performance systems, inference optimization, and model acceleration. You will thrive in ambiguity, tackle unclear problems, and design impactful solutions. The position offers a competitive base salary between $270,000 and $500,000, along with bonuses, equity, and benefits. Ideal candidates hold a relevant PhD or equivalent experience and have a knack for full-cycle ownership in engineering projects. Join a dynamic team prioritizing performance and collaboration.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Staff ML Engineer: Build Ultra-Fast AI at Scale (Relocation)
Staff ML Engineer: Build Ultra-Fast AI at Scale (Relocation)

Inworld AI • Mountain View (CA)

On-site
USD 270,000 - 500,000
Founding ML Inference Engineer — Ultra-Low Latency AI
Founding ML Inference Engineer — Ultra-Low Latency AI

Reactor • San Francisco (CA)

On-site
USD 180,000 - 280,000
Competitive salary
Early equity
Health, dental, and vision coverage
+1
Staff / Principal Machine Learning Engineer, Serving - USA
Staff / Principal Machine Learning Engineer, Serving - USA

Inworld • Mountain View (CA)

On-site
USD 270,000 - 500,000
Relocation assistance
Equity options
Comprehensive benefits
Staff / Principal Machine Learning Engineer, Serving
Staff / Principal Machine Learning Engineer, Serving

Inworld AI • Mountain View (CA)

On-site
USD 270,000 - 500,000
Relocation assistance
Equity options
Comprehensive benefits package
Software Engineer - ML Model Performance
Software Engineer - ML Model Performance

Baseten • San Francisco (CA)

On-site
USD 150,000 - 250,000
Staff ML Inference Engineer — Model Efficiency (Remote)
Staff ML Inference Engineer — Model Efficiency (Remote)

Jaide Health • San Francisco (CA)

On-site
USD 120,000 - 160,000
Inclusive culture and work environment
Weekly lunch stipend, in-office lunches & snacks
Full health and dental benefits
+2
Senior ML Performance Engineer — Ultra-Fast Inference + Equity
Senior ML Performance Engineer — Ultra-Fast Inference + Equity

well-funded deeptech startup • California (MO)

On-site
USD 200,000 - 250,000
Senior Inference Performance Engineer — GPU & CUDA
Senior Inference Performance Engineer — GPU & CUDA

Inference • San Francisco (CA)

Hybrid
USD 220,000 - 320,000
Competitive compensation
Equity in a high-growth startup
Comprehensive benefits
Staff Engineer - ML Inference & Model Efficiency
Staff Engineer - ML Inference & Model Efficiency

Cohere • San Francisco (CA)

Remote
USD 180,000 - 240,000
Inclusive work culture
Weekly lunch stipend
Full health and dental benefits
+4
ML Engineer - Inference & Model Deployment
ML Engineer - Inference & Model Deployment

HiringCafe • Cupertino (CA)

On-site
USD 250,000 - 310,000
Generous health, dental, and vision coverage
Paid parental leave
Relocation support