Senior Inference Systems Engineer — Large-Scale GPUs

RadixArk

Palo Alto (CA)

On-site

USD 190,000 - 260,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Competitive compensation
Meaningful equity
Comprehensive benefits
Flexible work arrangements

Job summary

A leading AI infrastructure company in California is seeking a Member of Technical Staff — Inference to design and optimize large-scale AI inference systems. The role demands 5+ years in systems engineering and expertise in large-scale inference systems. Successful candidates will enhance GPU utilization and work closely with various teams to debug and drive the reliability of infrastructure. Competitive compensation and flexible work arrangements are offered, alongside a commitment to equity and diversity.

Qualifications

  • 5+ years of experience in systems engineering, ML infrastructure, or performance-critical backend systems.
  • Strong expertise in large-scale inference systems for LLMs or generative models.
  • Deep understanding of GPU architecture and performance characteristics.
  • Experience optimizing latency- and throughput-critical production systems.
  • Strong knowledge of distributed systems and networking fundamentals.
  • Proficiency in C++, Rust, Go, or Python for production systems.
  • Experience profiling and optimizing compute-intensive workloads.

Responsibilities

  • Design and build large-scale inference systems for frontier AI models.
  • Optimize latency, throughput, and GPU utilization in production inference.
  • Develop and improve model serving architectures and runtimes.
  • Work on batching, scheduling, and memory management strategies.
  • Collaborate with kernel, compiler, and systems teams on performance optimization.
  • Debug performance bottlenecks across the stack.
  • Drive reliability and scalability of inference infrastructure.
  • Build tooling for observability, profiling, and performance analysis.
  • Contribute to long-term inference architecture and strategy.

Job description

A leading AI infrastructure company in California is seeking a Member of Technical Staff — Inference to design and optimize large-scale AI inference systems. The role demands 5+ years in systems engineering and expertise in large-scale inference systems. Successful candidates will enhance GPU utilization and work closely with various teams to debug and drive the reliability of infrastructure. Competitive compensation and flexible work arrangements are offered, alongside a commitment to equity and diversity.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Inference Performance Engineer — GPU & CUDA
Senior Inference Performance Engineer — GPU & CUDA

Inference • San Francisco (CA)

Hybrid
USD 220,000 - 320,000
Competitive compensation
Equity in a high-growth startup
Comprehensive benefits
Senior Inference Performance Engineer - GPU & CUDA
Senior Inference Performance Engineer - GPU & CUDA

inference.net • San Francisco (CA)

Hybrid
USD 220,000 - 320,000
Senior System Software Engineer - GPU AI Inference Equity
Senior System Software Engineer - GPU AI Inference Equity

NVIDIA • United States

On-site
USD 152,000 - 242,000
Senior GPU Systems Engineer: Large-Scale Inference & RL
Senior GPU Systems Engineer: Large-Scale Inference & RL

Reflection • New York (NY)

On-site
USD 150,000 - 200,000
Senior AI Inference Systems Engineer: GPU-Optimized, Cloud
Senior AI Inference Systems Engineer: GPU-Optimized, Cloud

NVIDIA Gruppe • Santa Clara (CA)

On-site
USD 184,000 - 356,500
Equity opportunities
Comprehensive benefits package
Senior GPU ML Infra Engineer — Mid-Training & Inference
Senior GPU ML Infra Engineer — Mid-Training & Inference

Reflection AI • San Francisco (CA)

On-site
USD 120,000 - 160,000
Senior AI Inference Performance Engineer (Remote)
Senior AI Inference Performance Engineer (Remote)

DigitalOcean • San Francisco (CA)

Remote
USD 167,000 - 209,000
GPU Performance Engineer: Scale AI Inference
GPU Performance Engineer: Scale AI Inference

Anthropic • San Francisco (CA)

On-site
USD 315,000 - 560,000
Competitive salary
Equity opportunities
Flexible working hours
+1
Senior Inference Platform Architect - GPU Cloud (Equity)
Senior Inference Platform Architect - GPU Cloud (Equity)

Hamilton Barnes Associates Limited • San Francisco (CA)

On-site
USD 270,000 - 330,000
Equity
Senior AI Inference Infrastructure Engineer
Senior AI Inference Infrastructure Engineer

Modular • United States

Hybrid
USD 167,000 - 273,000