Senior LLM Infra Engineer — HPC, Benchmarking & DevEx

Baseten

United States

Remote

USD 180,000 - 300,000

Full time

14 days+
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Job summary

Baseten is seeking a Senior Software Engineer to join a high-impact team at the crossroads of HPC and Large Language Model engineering. You will help define the roadmap, drive key technical decisions, and own the future of this work.

You will build automated benchmarks, GPU-enabled dev environments, and open-source tools like InferenceMAX and genai-bench, while profiling performance and optimizing the model runtimes stack.

Qualifications

  • Mid-senior level with high-leverage impact.
  • Strong communication skills to drive cross-team efforts.
  • Ability to navigate vague requirements and mentor other engineers.

Responsibilities

  • Benchmarking: Evaluate, run and automate LLM quality benchmarks and custom performance suites.
  • DevEx Improvement: Develop GPU-enabled development environments for model experimentation.
  • Tool Development: Build/open-source tools like InferenceMAX and genai-bench for evaluation and benchmarking.
  • System Profiling: Use profilers to collect performance profiles and identify bottlenecks.
  • Monitoring Observability: Develop real-time dashboards and alerts for system health and runtime performance.
  • Continuous Integration: Automate performance testing via CI/CD pipelines and release workflow automation.
  • Optimization Automation: Build tools to identify the Pareto frontier for latency, cost, and quality.

Skills

Systems hardware
Automation mindset
Mathematical curiosity
Strong communication
Mentoring
NVIDIA Nsight Systems
PyTorch Profiler
py-spy

Tools

NVIDIA Nsight Systems
PyTorch Profiler
py-spy

Job description

Baseten is seeking a Senior Software Engineer to join a high-impact team at the crossroads of HPC and Large Language Model engineering. You will help define the roadmap, drive key technical decisions, and own the future of this work.

You will build automated benchmarks, GPU-enabled dev environments, and open-source tools like InferenceMAX and genai-bench, while profiling performance and optimizing the model runtimes stack.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Software Engineer, LLM Performance Tooling
Senior Software Engineer, LLM Performance Tooling

BaseTen • San Francisco (CA), New York (NY)

On-site
USD 180,000 - 240,000
Equity
Health plans for dependents
Flexible PTO (Winter Break)
+3
LLM Performance Engineer — GPU & HPC
LLM Performance Engineer — GPU & HPC

Baseten • San Francisco (CA)

On-site
USD 160,000 - 200,000
Competitive compensation
Equity
Medical/dental/vision insurance
+4
Senior LLM Inference: GPU Kernel Optimization
Senior LLM Inference: GPU Kernel Optimization

NVIDIA • Austin (TX)

On-site
USD 184,000
Equity
Benefits
Software Engineer, Inference Stack (LLM Infra)
Software Engineer, Inference Stack (LLM Infra)

Baseten • United States

Remote
USD 180,000 - 240,000
Staff Software Engineer: LLM Inference & GPU Infra (Equity)
Staff Software Engineer: LLM Inference & GPU Infra (Equity)

DevHub • San Francisco (CA)

On-site
USD 180,000 - 240,000
Annual performance bonus
Equity
Staff Engineer - AI Inference & Benchmarking
Staff Engineer - AI Inference & Benchmarking

Liquid-Ai • Boston (MA)

Hybrid
USD 140,000 - 210,000
Equity
Health insurance
401(k) match
+1
Distributed LLM Inference Engineer - Scale & Resilience
Distributed LLM Inference Engineer - Scale & Resilience

Baseten • San Francisco (CA)

On-site
USD 180,000 - 360,000
Equity
Health coverage
Flexible PTO
+4
Senior LLM Efficiency Architect: Model Systems Co-Design
Senior LLM Efficiency Architect: Model Systems Co-Design

Nvidia Corporation in • Santa Clara (CA)

Hybrid
USD 184,000 - 357,000
Equity
Benefits
Staff Engineer - LLM Benchmark Platform
Staff Engineer - LLM Benchmark Platform

Vals AI, Inc. • San Francisco (CA)

On-site
USD 150,000 - 230,000
Relocation assistance
Health insurance
Dental insurance
+1
Inference Systems Engineer: Benchmarking & Porting (Hybrid)
Inference Systems Engineer: Benchmarking & Porting (Hybrid)

Liquid AI • Boston (MA)

Hybrid
USD 180,000 - 240,000
Equity
Health insurance
401(k) matching
+2