LLM Performance Engineer — GPU & HPC

Baseten

San Francisco (CA)

On-site

USD 160,000 - 200,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Competitive compensation
Equity
Medical/dental/vision insurance
Flexible PTO
Parental leave
401(k)
Learning opportunities

Job summary

Baseten is hiring early-career Software Engineers to join a team at the intersection of high-performance computing and large language model engineering. You will build automated speedometer and diagnostic tools for GPU clusters, measure FPS and memory bandwidth, and develop dev environments for model experimentation.

You will work across Linux, Python, and NVIDIA software to push performance and reliability in production-grade AI infrastructure.

Qualifications

  • Interest in HPC and LLM engineering at scale.
  • Willingness to learn and explore across hardware and software.
  • Familiarity with Python and eagerness to master NVIDIA stack.

Responsibilities

  • Run and automate LLM quality benchmarks and performance suites.
  • Create automated tests for GPU clusters across x86 and ARM.
  • Develop GPU-enabled dev environments optimized for model experimentation.
  • Build and contribute to tools to automate model evaluation and optimization.
  • Profile performance with PyTorch Profiler and Nsight Systems.
  • Develop real-time dashboards and CI/CD for performance tests.
  • Identify optimal configurations balancing latency, cost and quality.

Skills

Python
NVIDIA software stack
C++ (nice to have)
GPU architecture
Automation & scripting
Transformers / FLOPs concepts

Tools

PyTorch Profiler
NVIDIA Nsight Systems
genai-bench
InferenceMAX
CI/CD pipelines

Job description

Baseten is hiring early-career Software Engineers to join a team at the intersection of high-performance computing and large language model engineering. You will build automated speedometer and diagnostic tools for GPU clusters, measure FPS and memory bandwidth, and develop dev environments for model experimentation.

You will work across Linux, Python, and NVIDIA software to push performance and reliability in production-grade AI infrastructure.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

LLM Performance Engineer: GPU Benchmarks & Systems Tooling
LLM Performance Engineer: GPU Benchmarks & Systems Tooling

The Consensus • New York (NY)

On-site
USD 80,000 - 100,000
Competitive compensation
100% medical, dental, and vision coverage
Flexible PTO policy
+3
LLM Inference Performance Engineer - GPU Kernel Optimizer
LLM Inference Performance Engineer - GPU Kernel Optimizer

Intel • Folsom (CA)

Hybrid
USD 171,000 - 315,000
Stock bonuses
Health benefits
Hybrid work model
Software Engineer- Model Performance Systems
Software Engineer- Model Performance Systems

Baseten • San Francisco (CA)

On-site
USD 160,000 - 200,000
Competitive compensation
Equity
Medical/dental/vision insurance
+4
Senior LLM Infra Engineer — HPC, Benchmarking & DevEx
Senior LLM Infra Engineer — HPC, Benchmarking & DevEx

Baseten • United States

Remote
USD 180,000 - 300,000
GPU Inference Performance Engineer — Equity & Optimization
GPU Inference Performance Engineer — Equity & Optimization

Nvidia Corporation • Santa Clara (CA)

On-site
USD 152,000 - 242,000
Remote CUDA HPC Engineer—GPU Performance & ML Inference
Remote CUDA HPC Engineer—GPU Performance & ML Inference

Bright Vision Technologies • Flower Mound (TX)

On-site
USD 130,000 - 150,000
Remote work
Software Engineer, Model Performance Systems
Software Engineer, Model Performance Systems

Baseten • New York (NY)

On-site
USD 160,000 - 200,000
Competitive compensation with equity
100% medical, dental, and vision insurance coverage
Generous PTO including Winter Break
+3
Senior LLM Inference Engineer — Performance & GPU Optimization
Senior LLM Inference Engineer — Performance & GPU Optimization

Confidential • United States

On-site
USD 180,000 - 240,000
Lead GPU Performance Engineer — HPC Systems
Lead GPU Performance Engineer — HPC Systems

Nebius • United States

On-site
USD 170,000 - 300,000
Health insurance
401(k) plan
Parental leave
+2
Software Engineer, Model Performance Systems
Software Engineer, Model Performance Systems

The Consensus • New York (NY)

On-site
USD 80,000 - 100,000
Competitive compensation
100% medical, dental, and vision coverage
Flexible PTO policy
+3