Staff ML Performance Engineer — High-Throughput Inference
fal
San Francisco (CA)
On-site
USD 180,000 - 250,000
Full time
14 days+
Get more replies from employers
Send a job-specific resume in minutes.
Start fresh or import an existing resume
Benefits offered by this job
Competitive salary and equity
Visa sponsorship
Health, dental, and vision insurance
Regular team events
Job summary
A tech company in San Francisco is looking for a Staff Software Engineer to enhance model performance for generative media models. The ideal candidate will design advanced model serving architectures, develop performance tools, and collaborate with the Applied ML team. Offering a competitive salary of $180,000 - $250,000 with equity and comprehensive health benefits, this position provides substantial growth opportunities.
Qualifications
Strong foundation in systems programming with expertise in identifying and fixing bottlenecks.
Deep understanding of cutting edge ML infrastructure stack including model compilation, quantization, and serving architectures.
Proficient in Triton or willingness to learn with comparable experience in lower-level accelerator programming.
Responsibilities
Help maintain frontier position on model performance for generative media models.
Design and implement novel approaches to model serving architecture.
Develop performance monitoring and profiling tools.
Skills
Systems programming expertise
ML infrastructure stack knowledge
Multi-dimensional model parallelism
Triton proficiency
Tools
PyTorch
TensorRT
Nvidia systems
Job description
A tech company in San Francisco is looking for a Staff Software Engineer to enhance model performance for generative media models. The ideal candidate will design advanced model serving architectures, develop performance tools, and collaborate with the Applied ML team. Offering a competitive salary of $180,000 - $250,000 with equity and comprehensive health benefits, this position provides substantial growth opportunities.