Senior Software Engineer, LLM Performance Tooling

BaseTen

San Francisco, New York (CA, NY)

On-site

USD 180,000 - 240,000

Full time

14 days+
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Benefits offered by this job

Equity
Health plans for dependents
Flexible PTO (Winter Break)
Parental leave
Fertility stipend
401(k)

Job summary

Baseten is seeking a mid-to-senior Software Engineer to advance HPC and LLM infrastructure. You will define roadmaps, own critical tooling, and drive performance optimizations across GPU clusters for production-scale AI workloads.

You will build internal development environments, automate benchmarking, and contribute to open-source projects, shaping Baseten's platform for cutting-edge AI deployments.

Qualifications

  • A love for systems & hardware: GPU memory subsystems, InfiniBand, data movement in clusters.
  • An automation mindset: scripting repetitive tasks, stress-testing, and breaking points.
  • Mathematical curiosity: understanding Transformers, FLOPs, memory requirements.
  • Technical toolkit: proficiency with Python and NVIDIA software stack; knowledge of C++ is a plus.

Responsibilities

  • Benchmarking: run and automate LLM quality benchmarks (GSM8K, MMLU) and workload-specific tests.
  • DevEx improvements: build GPU-enabled development environments for fast model experimentation.
  • Tool development: contribute to open-source tools for evaluation and benchmarking.
  • System profiling: use profilers to identify bottlenecks in compute and networking stacks.
  • Monitoring & observability: create real-time dashboards and alerts for system health and startup times.
  • CI: automate performance tests and release workflows for the runtime stack.
  • Optimization automation: discover Pareto-optimal configurations for latency, cost, and quality.

Skills

Python
NVIDIA software stack
C++ familiarity

Job description

Baseten is seeking a mid-to-senior Software Engineer to advance HPC and LLM infrastructure. You will define roadmaps, own critical tooling, and drive performance optimizations across GPU clusters for production-scale AI workloads.

You will build internal development environments, automate benchmarking, and contribute to open-source projects, shaping Baseten's platform for cutting-edge AI deployments.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

LLM Performance Engineer — GPU & HPC
LLM Performance Engineer — GPU & HPC

Baseten • San Francisco (CA)

On-site
USD 160,000 - 200,000
Competitive compensation
Equity
Medical/dental/vision insurance
+4
Senior LLM Infra Engineer — HPC, Benchmarking & DevEx
Senior LLM Infra Engineer — HPC, Benchmarking & DevEx

Baseten • United States

Remote
USD 180,000 - 300,000
Software Engineer, Model Performance Tooling
Software Engineer, Model Performance Tooling

BaseTen • San Francisco (CA), New York (NY)

On-site
USD 180,000 - 240,000
Equity
Health plans for dependents
Flexible PTO (Winter Break)
+3
Software Engineer- Model Performance Systems
Software Engineer- Model Performance Systems

Baseten • San Francisco (CA)

On-site
USD 160,000 - 200,000
Competitive compensation
Equity
Medical/dental/vision insurance
+4
Software Engineer, Model Performance Systems
Software Engineer, Model Performance Systems

The Consensus • New York (NY)

On-site
USD 80,000 - 100,000
Competitive compensation
100% medical, dental, and vision coverage
Flexible PTO policy
+3
Engineering Manager, Forward Deployed AI & LLM Inference
Engineering Manager, Forward Deployed AI & LLM Inference

BaseTen • New York (NY), San Francisco (CA)

On-site
USD 180,000 - 240,000
Competitive equity
Medical, dental, vision coverage
Flexible PTO
+3
Software Engineer - Model Performance
Software Engineer - Model Performance

Baseten • San Francisco (CA)

On-site
USD 150,000 - 250,000
Competitive compensation with equity
100% medical, dental, and vision insurance
Generous PTO policy
+2
Software Engineer - Model Performance
Software Engineer - Model Performance

Baseten • New York (NY)

On-site
USD 200,000 - 275,000
100% coverage of medical, dental, and vision insurance
Generous PTO policy
Paid parental leave
+2
LLM Performance Engineer: GPU Benchmarks & Systems Tooling
LLM Performance Engineer: GPU Benchmarks & Systems Tooling

The Consensus • New York (NY)

On-site
USD 80,000 - 100,000
Competitive compensation
100% medical, dental, and vision coverage
Flexible PTO policy
+3
AI Inference Platform Engineer
AI Inference Platform Engineer

BaseTen • New York (NY), San Francisco (CA)

On-site
USD 140,000 - 210,000
Equity
Medical, dental and vision insurance (
Flexible PTO including Winter Break
+4