Staff Software Engineer, ML Performance & Systems

fal

San Francisco (CA)

On-site

USD 180,000 - 250,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Competitive salary and equity
Visa sponsorship
Health, dental, and vision insurance
Regular team events

Job summary

A tech company in San Francisco is looking for a Staff Software Engineer to enhance model performance for generative media models. The ideal candidate will design advanced model serving architectures, develop performance tools, and collaborate with the Applied ML team. Offering a competitive salary of $180,000 - $250,000 with equity and comprehensive health benefits, this position provides substantial growth opportunities.

Qualifications

  • Strong foundation in systems programming with expertise in identifying and fixing bottlenecks.
  • Deep understanding of cutting edge ML infrastructure stack including model compilation, quantization, and serving architectures.
  • Proficient in Triton or willingness to learn with comparable experience in lower-level accelerator programming.

Responsibilities

  • Help maintain frontier position on model performance for generative media models.
  • Design and implement novel approaches to model serving architecture.
  • Develop performance monitoring and profiling tools.

Skills

Systems programming expertise
ML infrastructure stack knowledge
Multi-dimensional model parallelism
Triton proficiency

Tools

PyTorch
TensorRT
Nvidia systems

Job description

Staff Software Engineer, ML Performance & Systems

Help fal maintain its frontier position on model performance for generative media models. Design and implement novel approaches to model serving architecture on top of our in-house inference engine, focusing on maximizing throughput while minimizing latency and resource usage. Develop performance monitoring and profiling tools to identify bottlenecks and optimization opportunities. Work closely with our Applied ML team and customers (frontier labs on the media space) and make sure their workloads benefit from our accelerator.

Key Responsibilities
  • Help fal maintain its frontier position on model performance for generative media models.
  • Design and implement novel approaches to model serving architecture on top of our in-house inference engine, focusing on maximizing throughput while minimizing latency and resource usage.
  • Develop performance monitoring and profiling tools to identify bottlenecks and optimization opportunities.
  • Work closely with our Applied ML team and customers (frontier labs on the media space) and make sure their workloads benefit from our accelerator.
Requirements
  • Strong foundation in systems programming with expertise in identifying and fixing bottlenecks.
  • Deep understanding of cutting edge ML infrastructure stack (anything from PyTorch, TensorRT, TransformerEngine to Nsight), including model compilation, quantization, and serving architectures. Ideally following closely the developments in all these systems as they happen.
  • Have a fundamental view of the underlying hardware (Nvidia based systems at the moment), and when necessary go deeper into the stack to fix bottlenecks (custom GEMM kernels with CUTLASS for common shapes).
  • Proficient in Triton or willingness to learn with comparable experience in lower-level accelerator programming.
  • New frontier: multi-dimensional model parallelism (combining multiple parallelism techniques like TP with context parallel / sequence parallel).
  • Familiar with internals of Ring Attention, FA3, FusedMLP implementations.
What we offer at fal
  • Interesting and challenging work
  • Competitive salary and equity
  • Employee-friendly equity terms (early exercise, extended exercise)
  • A lot of learning and growth opportunities
  • We offer visa sponsorship and will help you relocate to San Francisco.
  • Health, dental, and vision insurance (US)
  • Regular team events and offsite
Compensation

$180,000 - $250,000 + equity + comprehensive benefits package

Location

We are currently hiring in downtown San Francisco.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Software Engineer, Distributed Systems
Software Engineer, Distributed Systems

Fal • San Francisco (CA)

On-site
USD 180,000 - 250,000
Remote work options for senior levels
Visa sponsorship and relocation
Health, dental, and vision insurance
+1
Senior Software Engineer, Product
Senior Software Engineer, Product

The Consensus • San Francisco (CA)

On-site
USD 180,000 - 230,000
Relocation assistance to SF
Health, dental, and vision insurance (
Regular team events and offsite
Software Engineer, Infrastructure
Software Engineer, Infrastructure

The Consensus • San Francisco (CA)

Hybrid
USD 180,000 - 250,000
Relocation assistance
Health, dental, and vision insurance (
Team events & offsites
+1
Senior Software Engineer, Product
Senior Software Engineer, Product

Fal • San Francisco (CA)

Hybrid
USD 180,000 - 230,000
Relocation assistance
Health, dental, and vision insurance
Regular team events
+2
Senior Software Engineer, Core Product Systems
Senior Software Engineer, Core Product Systems

fal • San Francisco (CA)

On-site
USD 190,000 - 230,000
Relocation assistance + visa Spons. to
Health, dental, and vision insurance
Regular team events and offsites
Founding Engineer, ML Inference
Founding Engineer, ML Inference

Reactor • San Francisco (CA)

On-site
USD 180,000 - 280,000
Competitive salary
Early equity
Health, dental, and vision coverage
+1
Software Engineer, Growth
Software Engineer, Growth

fal • San Francisco (CA)

On-site
USD 170,000 - 220,000
Relocation assistance to San Francisco
Health, dental, and vision insurance (
Team events & offsites
+2
Software Engineer, Site Reliability
Software Engineer, Site Reliability

Fal • San Francisco (CA)

On-site
USD 180,000 - 250,000
Relocation assistance to San Francisco
Visa sponsorship
Competitive salary and equity
+1
Staff Technical Lead for Inference & ML Performance
Staff Technical Lead for Inference & ML Performance

fal • San Francisco (CA)

On-site
USD 150,000 - 200,000
Software Engineer, Infrastructure
Software Engineer, Infrastructure

fal • San Francisco (CA)

On-site
USD 180,000 - 250,000
Relocation assistance to San Francisco
Health, dental and vision insurance (U
Regular team events and offsites