Staff ML Engineer: Efficient ML & Low-Latency AI

Embedding VC

San Francisco (CA)

On-site

USD 100,000 - 150,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

A tech-focused company in San Francisco seeks candidates with expertise in AI simulation development. The role emphasizes optimizing training efficiency, enhancing GPU performance, and ensuring low-latency inference. Applicants should be proficient in methodologies for gradient checkpointing, Nsight profiling, and job management tools like SLURM. The company values in-person collaboration in its dynamic team environment, providing opportunities for innovation and cutting-edge technology implementation.

Responsibilities

  • Optimize training efficiency with ML techniques.
  • Enhance GPU and kernel performance for AI models.
  • Implement inference optimization strategies for low-latency serving.
  • Ensure infra reliability with job management tools.

Skills

Dataloaders, fusion, activation remat
Gradient checkpointing
Nsight profiling
Triton/CUDA kernels
Quantization (GPTQ/AWQ)

Tools

SLURM
Kubernetes

Job description

A tech-focused company in San Francisco seeks candidates with expertise in AI simulation development. The role emphasizes optimizing training efficiency, enhancing GPU performance, and ensuring low-latency inference. Applicants should be proficient in methodologies for gradient checkpointing, Nsight profiling, and job management tools like SLURM. The company values in-person collaboration in its dynamic team environment, providing opportunities for innovation and cutting-edge technology implementation.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior ML Engineer - Low-Latency Inference & Systems
Senior ML Engineer - Low-Latency Inference & Systems

Inworld • Germany (OH)

Hybrid
USD 120,000 - 160,000
Staff ML Infrastructure & Performance Engineer
Staff ML Infrastructure & Performance Engineer

Embedding VC • San Mateo (CA)

On-site
USD 120,000 - 150,000
Staff ML Performance Engineer — Scalable Inference & CUDA
Staff ML Performance Engineer — Scalable Inference & CUDA

Modal • New York (NY)

On-site
USD 120,000 - 160,000
Software Engineer - ML Model Performance
Software Engineer - ML Model Performance

Baseten • San Francisco (CA)

On-site
USD 150,000 - 250,000
Staff ML Engineer: Build Ultra-Fast AI at Scale (Relocation)
Staff ML Engineer: Build Ultra-Fast AI at Scale (Relocation)

Inworld AI • Mountain View (CA)

On-site
USD 270,000 - 500,000
Senior ML Infra Engineer — Real-Time, Low-Latency
Senior ML Infra Engineer — Real-Time, Low-Latency

LMArena • San Francisco (CA)

On-site
USD 170,000 - 260,000
Competitive compensation
Comprehensive health benefits
Opportunity to work on cutting-edge AI
Founding ML Inference Engineer — Ultra-Low Latency AI
Founding ML Inference Engineer — Ultra-Low Latency AI

Reactor • San Francisco (CA)

On-site
USD 180,000 - 280,000
Competitive salary
Early equity
Health, dental, and vision coverage
+1
Senior ML Engineer - Real-Time Inference & Systems
Senior ML Engineer - Real-Time Inference & Systems

Inworld AI • Germany (OH)

On-site
USD 120,000 - 180,000
Senior GPU ML Infra Engineer — Mid-Training & Inference
Senior GPU ML Infra Engineer — Mid-Training & Inference

Reflection AI • San Francisco (CA)

On-site
USD 120,000 - 160,000
Staff Engineer - ML Inference & Model Efficiency
Staff Engineer - ML Inference & Model Efficiency

Cohere • San Francisco (CA)

Remote
USD 180,000 - 240,000
Inclusive work culture
Weekly lunch stipend
Full health and dental benefits
+4