Senior Inference Performance Engineer - GPU & CUDA

inference.net

San Francisco (CA)

Hybrid

USD 220,000 - 320,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Equity in a high-growth startup
Comprehensive benefits

Job summary

inference.net, a growing company in San Francisco, seeks an experienced engineer to optimize AI inference performance. The ideal candidate will have over 2 years of experience in ML systems and GPU programming. Key responsibilities include implementing optimization techniques and debugging inference frameworks. The role offers a competitive salary of $220,000 - $320,000 plus equity, and values curiosity and fast learning. You will join a team devoted to innovative AI solutions.

Qualifications

  • 2+ years of experience in ML systems, inference optimization, or GPU programming.
  • Strong proficiency in Python and familiarity with C++.
  • Hands-on experience with LLM inference frameworks.

Responsibilities

  • Implement and productionize optimization techniques.
  • Deep dive into inference frameworks to debug and improve performance.
  • Profile and optimize CUDA kernels and GPU utilization.

Skills

Machine Learning systems
Inference optimization
GPU programming
Python
C++
LLM inference frameworks
GPU architecture
Performance improvement

Tools

PyTorch
CUDA
Docker
Kubernetes

Job description

inference.net, a growing company in San Francisco, seeks an experienced engineer to optimize AI inference performance. The ideal candidate will have over 2 years of experience in ML systems and GPU programming. Key responsibilities include implementing optimization techniques and debugging inference frameworks. The role offers a competitive salary of $220,000 - $320,000 plus equity, and values curiosity and fast learning. You will join a team devoted to innovative AI solutions.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Inference Performance Engineer — GPU & CUDA
Senior Inference Performance Engineer — GPU & CUDA

Inference • San Francisco (CA)

Hybrid
USD 220,000 - 320,000
Competitive compensation
Equity in a high-growth startup
Comprehensive benefits
GPU Performance Engineer: Scale AI Inference
GPU Performance Engineer: Scale AI Inference

Anthropic • San Francisco (CA)

On-site
USD 315,000 - 560,000
Competitive salary
Equity opportunities
Flexible working hours
+1
Senior Inference Performance Engineer — Equity & Hybrid
Senior Inference Performance Engineer — Equity & Hybrid

NVIDIA Gruppe • Santa Clara (CA)

Hybrid
USD 124,000 - 242,000
Senior AI Inference Engineer - GPU, Rust & CUDA
Senior AI Inference Engineer - GPU, Rust & CUDA

Perplexity • San Francisco (CA)

On-site
USD 220,000 - 485,000
Senior Inference Systems Engineer — Large-Scale GPUs
Senior Inference Systems Engineer — Large-Scale GPUs

RadixArk • Palo Alto (CA)

On-site
USD 190,000 - 260,000
Competitive compensation
Meaningful equity
Comprehensive benefits
+1
Senior GPU Inference Engine Engineer
Senior GPU Inference Engine Engineer

FriendliAI • San Francisco (CA)

On-site
USD 120,000 - 160,000
Flexible working hours
Daily lunch and dinner provided; unlimited snacks and beverages
Health check-up support and top-tier equipment/hardware support
+2
Senior AI Inference Performance Engineer (Remote)
Senior AI Inference Performance Engineer (Remote)

DigitalOcean • San Francisco (CA)

Remote
USD 167,000 - 209,000
Senior AI Inference Performance Architect
Senior AI Inference Performance Architect

NVIDIA • California (MO)

On-site
USD 152,000 - 242,000
Senior GPU ML Infra Engineer — Mid-Training & Inference
Senior GPU ML Infra Engineer — Mid-Training & Inference

Reflection AI • San Francisco (CA)

On-site
USD 120,000 - 160,000
Inference Performance Engineer: AI GPU Optimization&Equity
Inference Performance Engineer: AI GPU Optimization&Equity

NVIDIA • Santa Clara (CA)

Hybrid
USD 124,000 - 242,000
Equity
Benefits package