LLM Performance Engineer — GPU & HPC

Baseten

San Francisco (CA)

On-site

USD 160,000 - 200,000

Full time

14 days+
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Benefits offered by this job

Competitive compensation
Equity
Medical/dental/vision insurance
Flexible PTO
Parental leave
401(k)
Learning opportunities

Job summary

Baseten is hiring early-career Software Engineers to join a team at the intersection of high-performance computing and large language model engineering. You will build automated speedometer and diagnostic tools for GPU clusters, measure FPS and memory bandwidth, and develop dev environments for model experimentation.

You will work across Linux, Python, and NVIDIA software to push performance and reliability in production-grade AI infrastructure.

Qualifications

  • Interest in HPC and LLM engineering at scale.
  • Willingness to learn and explore across hardware and software.
  • Familiarity with Python and eagerness to master NVIDIA stack.

Responsibilities

  • Run and automate LLM quality benchmarks and performance suites.
  • Create automated tests for GPU clusters across x86 and ARM.
  • Develop GPU-enabled dev environments optimized for model experimentation.
  • Build and contribute to tools to automate model evaluation and optimization.
  • Profile performance with PyTorch Profiler and Nsight Systems.
  • Develop real-time dashboards and CI/CD for performance tests.
  • Identify optimal configurations balancing latency, cost and quality.

Skills

Python
NVIDIA software stack
C++ (nice to have)
GPU architecture
Automation & scripting
Transformers / FLOPs concepts

Tools

PyTorch Profiler
NVIDIA Nsight Systems
genai-bench
InferenceMAX
CI/CD pipelines

Job description

Baseten is hiring early-career Software Engineers to join a team at the intersection of high-performance computing and large language model engineering. You will build automated speedometer and diagnostic tools for GPU clusters, measure FPS and memory bandwidth, and develop dev environments for model experimentation.

You will work across Linux, Python, and NVIDIA software to push performance and reliability in production-grade AI infrastructure.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior Software Engineer, LLM Performance Tooling
Senior Software Engineer, LLM Performance Tooling

BaseTen • San Francisco (CA), New York (NY)

On-site
USD 180,000 - 240,000
Equity
Health plans for dependents
Flexible PTO (Winter Break)
+3
LLM Performance Engineer: GPU Benchmarks & Systems Tooling
LLM Performance Engineer: GPU Benchmarks & Systems Tooling

The Consensus • New York (NY)

On-site
USD 80,000 - 100,000
Competitive compensation
100% medical, dental, and vision coverage
Flexible PTO policy
+3
LLM Inference Performance Engineer
LLM Inference Performance Engineer

Xapply • San Francisco (CA)

On-site
USD 170,000 - 230,000
Competitive compensation and equity
Medical, dental, vision insurance (US)
Flexible PTO including Winter Break
+4
Performance Benchmark Engineer (Equity) for AI/LLM
Performance Benchmark Engineer (Equity) for AI/LLM

NVIDIA Corporation • Santa Clara (CA)

On-site
USD 136,000 - 213,000
Equity
Benefits
Senior LLM Training Performance Engineer (Hybrid)
Senior LLM Training Performance Engineer (Hybrid)

NVIDIA Corporation • Santa Clara (CA)

Hybrid
USD 184,000 - 357,000
Equity
Benefits package
Hybrid work environment
Senior LLM Infra Engineer — HPC, Benchmarking & DevEx
Senior LLM Infra Engineer — HPC, Benchmarking & DevEx

Baseten • United States

Remote
USD 180,000 - 300,000
Software Engineer- Model Performance Systems
Software Engineer- Model Performance Systems

Baseten • San Francisco (CA)

On-site
USD 160,000 - 200,000
Competitive compensation
Equity
Medical/dental/vision insurance
+4
LLM Inference Performance Engineer
LLM Inference Performance Engineer

Emploive • New York (NY)

On-site
USD 180,000 - 360,000
Equity
Insurance for dependents
Winter Break
Software Engineer, Model Performance Systems
Software Engineer, Model Performance Systems

The Consensus • New York (NY)

On-site
USD 80,000 - 100,000
Competitive compensation
100% medical, dental, and vision coverage
Flexible PTO policy
+3
Senior Performance Engineer: AI/LLM Benchmark Lead
Senior Performance Engineer: AI/LLM Benchmark Lead

NVIDIA Corporation • Santa Clara (CA)

Hybrid
USD 136,000 - 270,000
Equity compensation
Benefits