Staff ML Infrastructure & Performance Engineer

Embedding VC

San Mateo (CA)

On-site

USD 120,000 - 150,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

A leading AI company in San Mateo is seeking a candidate to optimize AI models for throughput, latency, and cost. The role involves working with advanced GPU performance techniques and a serving stack, focusing on delivering models at a faster and more cost-effective rate without compromising quality. Ideal candidates will have experience in GPU performance and be ready to join an on-site team. Strong familiarity with AI deployment systems is essential.

Qualifications

  • Experience in GPU performance and optimization.
  • Familiarity with AI model deployment and serving stack.
  • Strong background in parallel computing techniques.

Responsibilities

  • Optimize models for performance and cost.
  • Work with GPU performance and quantization.
  • Implement systems for observability and scaling.

Skills

CUDA/Triton kernels
Parallelism: FSDP/ZeRO
Quantization/PEFT
TensorRT-LLM/Triton Inference Server
Observability tools (Prom/Grafana/OpenTelemetry)

Job description

A leading AI company in San Mateo is seeking a candidate to optimize AI models for throughput, latency, and cost. The role involves working with advanced GPU performance techniques and a serving stack, focusing on delivering models at a faster and more cost-effective rate without compromising quality. Ideal candidates will have experience in GPU performance and be ready to join an on-site team. Strong familiarity with AI deployment systems is essential.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Staff ML Engineer: Efficient ML & Low-Latency AI
Staff ML Engineer: Efficient ML & Low-Latency AI

Embedding VC • San Francisco (CA)

On-site
USD 100,000 - 150,000
AI/ML Engineer: Next‑Gen Platforms & GPU Workloads
AI/ML Engineer: Next‑Gen Platforms & GPU Workloads

VeeAR Projects Inc. • Sunnyvale (CA)

On-site
USD 140,000 - 210,000
Staff ML Engineer: Build Ultra-Fast AI at Scale (Relocation)
Staff ML Engineer: Build Ultra-Fast AI at Scale (Relocation)

Inworld AI • Mountain View (CA)

On-site
USD 270,000 - 500,000
Software Engineer - ML Model Performance
Software Engineer - ML Model Performance

Baseten • San Francisco (CA)

On-site
USD 150,000 - 250,000
Senior ML Engineer - Low-Latency Inference & Systems
Senior ML Engineer - Low-Latency Inference & Systems

Inworld • Germany (OH)

Hybrid
USD 120,000 - 160,000
Staff ML Performance Engineer: Scale Training Throughput
Staff ML Performance Engineer: Scale Training Throughput

Wayve • Sunnyvale (CA)

On-site
USD 130,000 - 160,000
Senior Model Serving Engineer - Low-Latency AI Platform
Senior Model Serving Engineer - Low-Latency AI Platform

Menlo Ventures • San Francisco (CA)

On-site
USD 192,000 - 260,000
Comprehensive benefits
Eligibility for annual performance bonus
Equity opportunities
Fellow GPU Performance Optimizer for AI Training
Fellow GPU Performance Optimizer for AI Training

Advanced Micro Devices • San Jose (CA)

On-site
USD 140,000 - 180,000
Health insurance
Retirement plan
Paid time off
AI Infrastructure Performance Engineer
AI Infrastructure Performance Engineer

OpenAI • San Francisco (CA)

Hybrid
USD 120,000 - 160,000
Relocation assistance
Senior ML Performance Engineer: LLM Benchmarking & GPU
Senior ML Performance Engineer: LLM Benchmarking & GPU

Amadeus Search • San Francisco (CA)

Hybrid
USD 120,000 - 160,000
Competitive salary
Equity and bonus opportunities
Medical, dental, and vision coverage
+2