Staff ML Performance Engineer — High-Throughput Inference

fal

San Francisco (CA)

On-site

USD 180,000 - 250,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Competitive salary and equity
Visa sponsorship
Health, dental, and vision insurance
Regular team events

Job summary

A tech company in San Francisco is looking for a Staff Software Engineer to enhance model performance for generative media models. The ideal candidate will design advanced model serving architectures, develop performance tools, and collaborate with the Applied ML team. Offering a competitive salary of $180,000 - $250,000 with equity and comprehensive health benefits, this position provides substantial growth opportunities.

Qualifications

  • Strong foundation in systems programming with expertise in identifying and fixing bottlenecks.
  • Deep understanding of cutting edge ML infrastructure stack including model compilation, quantization, and serving architectures.
  • Proficient in Triton or willingness to learn with comparable experience in lower-level accelerator programming.

Responsibilities

  • Help maintain frontier position on model performance for generative media models.
  • Design and implement novel approaches to model serving architecture.
  • Develop performance monitoring and profiling tools.

Skills

Systems programming expertise
ML infrastructure stack knowledge
Multi-dimensional model parallelism
Triton proficiency

Tools

PyTorch
TensorRT
Nvidia systems

Job description

A tech company in San Francisco is looking for a Staff Software Engineer to enhance model performance for generative media models. The ideal candidate will design advanced model serving architectures, develop performance tools, and collaborate with the Applied ML team. Offering a competitive salary of $180,000 - $250,000 with equity and comprehensive health benefits, this position provides substantial growth opportunities.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Software Engineer - ML Model Performance
Software Engineer - ML Model Performance

Baseten • San Francisco (CA)

On-site
USD 150,000 - 250,000
Senior Staff Tech Lead — Inference & ML Performance
Senior Staff Tech Lead — Inference & ML Performance

fal • San Francisco (CA)

On-site
USD 150,000 - 200,000
Staff ML Infrastructure & Performance Engineer
Staff ML Infrastructure & Performance Engineer

Embedding VC • San Mateo (CA)

On-site
USD 120,000 - 150,000
Engineering Manager, ML Inference & Scale
Engineering Manager, ML Inference & Scale

Anthropic • San Francisco (CA)

Hybrid
USD 425,000 - 560,000
Competitive compensation
Flexible working hours
Generous vacation and parental leave
Founding ML Inference Engineer — Ultra-Low Latency AI
Founding ML Inference Engineer — Ultra-Low Latency AI

Reactor • San Francisco (CA)

On-site
USD 60,000 - 80,000
Staff ML Engineer: Build Ultra-Fast AI at Scale (Relocation)
Staff ML Engineer: Build Ultra-Fast AI at Scale (Relocation)

Inworld AI • Mountain View (CA)

On-site
USD 270,000 - 500,000
Senior Model Inference Engineer for Production-Scale AI
Senior Model Inference Engineer for Production-Scale AI

OpenAI • San Francisco (CA)

On-site
USD 325,000 - 490,000
Staff ML Performance Engineer: Scale Training Throughput
Staff ML Performance Engineer: Scale Training Throughput

Wayve • Sunnyvale (CA)

On-site
USD 130,000 - 160,000
Staff ML Engineer: Generative AI & Agentic Systems(Equity)
Staff ML Engineer: Generative AI & Agentic Systems(Equity)

Airbnb • San Francisco (CA)

Hybrid
USD 212,000 - 276,500
Bonus
Equity
Employee Travel Credits
Senior ML Engineer — Generative AI for Brand Content
Senior ML Engineer — Generative AI for Brand Content

Arcade • San Francisco (CA)

Hybrid
USD 180,000 - 300,000
Unlimited PTO
Rich health/401(k) plans
Meeting-light culture
+2