Member of Technical Staff — Inference

RadixArk

Palo Alto (CA)

On-site

USD 190,000 - 260,000

Full time

14 days+
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Benefits offered by this job

Competitive compensation
Meaningful equity
Comprehensive benefits
Flexible work arrangements

Job summary

A leading AI infrastructure company in California is seeking a Member of Technical Staff — Inference to design and optimize large-scale AI inference systems. The role demands 5+ years in systems engineering and expertise in large-scale inference systems. Successful candidates will enhance GPU utilization and work closely with various teams to debug and drive the reliability of infrastructure. Competitive compensation and flexible work arrangements are offered, alongside a commitment to equity and diversity.

Qualifications

  • 5+ years of experience in systems engineering, ML infrastructure, or performance-critical backend systems.
  • Strong expertise in large-scale inference systems for LLMs or generative models.
  • Deep understanding of GPU architecture and performance characteristics.
  • Experience optimizing latency- and throughput-critical production systems.
  • Strong knowledge of distributed systems and networking fundamentals.
  • Proficiency in C++, Rust, Go, or Python for production systems.
  • Experience profiling and optimizing compute-intensive workloads.

Responsibilities

  • Design and build large-scale inference systems for frontier AI models.
  • Optimize latency, throughput, and GPU utilization in production inference.
  • Develop and improve model serving architectures and runtimes.
  • Work on batching, scheduling, and memory management strategies.
  • Collaborate with kernel, compiler, and systems teams on performance optimization.
  • Debug performance bottlenecks across the stack.
  • Drive reliability and scalability of inference infrastructure.
  • Build tooling for observability, profiling, and performance analysis.
  • Contribute to long-term inference architecture and strategy.

Skills

Systems engineering
ML infrastructure
Performance optimization
Large-scale inference
GPU architecture
Distributed systems
C++, Python, Rust, Go
Profiling
Latency optimization

Tools

SGLang
CUDA
TensorRT-LLM
Kernels optimization

Job description

RadixArk is seeking a Member of Technical Staff — Inference to push the limits of large-scale AI inference.

You will work on the core systems that serve frontier models at scale, optimizing performance, latency, throughput, and cost across thousands of GPUs. This role sits at the intersection of systems engineering, ML infrastructure, and performance optimization.

Your work will directly shape how state‑of‑the‑art models are deployed and experienced by users worldwide.

This is a deeply technical, high-impact role for engineers who enjoy working close to the hardware–software boundary and solving performance‑critical problems at scale.

Requirements

5+ years of experience in systems engineering, ML infrastructure, or performance‑critical backend systems

Strong expertise in large‑scale inference systems for LLMs or generative models

Deep understanding of GPU architecture and performance characteristics

Experience optimizing latency- and throughput‑critical production systems

Strong knowledge of distributed systems and networking fundamentals

Proficiency in C++, Rust, Go, or Python for production systems

Experience profiling and optimizing compute‑intensive workloads

Strong Plus

Experience with LLM serving stacks (vLLM, TensorRT‑LLM, SGLang, etc.)

Familiarity with CUDA, Triton, or custom kernel optimization

Experience with batching, KV‑cache management, and scheduling strategies

Experience running inference at scale (1000+ GPUs)

Background in HPC or high‑performance systems

Open‑source contributions in ML or systems infrastructure

Responsibilities

Design and build large‑scale inference systems for frontier AI models

Optimize latency, throughput, and GPU utilization in production inference

Develop and improve model serving architectures and runtimes

Work on batching, scheduling, and memory management strategies

Collaborate with kernel, compiler, and systems teams on performance optimization

Debug performance bottlenecks across the stack

Drive reliability and scalability of inference infrastructure

Build tooling for observability, profiling, and performance analysis

Contribute to long‑term inference architecture and strategy

About RadixArk

RadixArk is an infrastructure‑first company built by engineers who've shipped production AI systems, created SGLang (20K+ GitHub stars, the fastest open LLM serving engine), and developed Miles (our large‑scale RL framework).

We're on a mission to democratize frontier‑level AI infrastructure by building world‑class open systems for inference and training.

Our team has optimized kernels serving billions of tokens daily and designed distributed systems coordinating 10,000+ GPUs across training and serving.

We're backed by leading infrastructure investors and collaborate with frontier AI labs and cloud providers.

Join us in building the infrastructure layer that powers the next generation of AI.

Compensation

We offer competitive compensation with meaningful equity, comprehensive benefits, and flexible work arrangements. Compensation depends on location, experience, and level.

RadixArk is an Equal Opportunity Employer and welcomes candidates from all backgrounds.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Member of Technical Staff — Inference-Kernel, Compiler & Communication
Member of Technical Staff — Inference-Kernel, Compiler & Communication

RadixArk • Palo Alto (CA)

On-site
USD 210,000 - 290,000
Competitive compensation
Comprehensive benefits
Flexible work arrangements
Member of Technical Staff — Cluster / Platform
Member of Technical Staff — Cluster / Platform

RadixArk • Palo Alto (CA)

On-site
USD 200,000 - 400,000
Member of Technical Staff — Training
Member of Technical Staff — Training

RadixArk • Palo Alto (CA)

On-site
USD 120,000 - 160,000
Comprehensive benefits
Flexible work arrangements
Member of Technical Staff — Heterogenous Hardware
Member of Technical Staff — Heterogenous Hardware

RadixArk • Palo Alto (CA)

On-site
USD 180,000 - 240,000
Competitive compensation
Equity
Flexible work arrangements
+1
AI Inference Engineer
AI Inference Engineer

Premier Global Links • Palo Alto (CA)

On-site
USD 230,000 - 350,000
Member of Technical Staff — Inference-Multimodal & Diffusion
Member of Technical Staff — Inference-Multimodal & Diffusion

RadixArk • Palo Alto (CA)

On-site
USD 180,000 - 280,000
Competitive compensation
Meaningful equity
Comprehensive benefits
+1
AI Inference Engineer
AI Inference Engineer

Premier Global Links LLC • Palo Alto (CA), Northern (KY)

Hybrid
USD 230,000 - 350,000
Equity 0.5%
AI Inference Engineer
AI Inference Engineer

Socket.dev • Palo Alto (CA)

On-site
USD 230,000 - 350,000
Equity opportunity (0.5%)
Professional growth
High-impact work
AI Infra Resident (1-Year Program)
AI Infra Resident (1-Year Program)

RadixArk • Palo Alto (CA)

On-site
USD 70,000 - 90,000
Health benefits
Potential for full-time position with equity
Member of Technical Staff, Inference Systems
Member of Technical Staff, Inference Systems

Confidential • California (MO)

On-site
USD 150,000 - 210,000