Member of Technical Staff — Inference-Kernel, Compiler & Communication

RadixArk

Palo Alto (CA)

On-site

USD 210,000 - 290,000

Full time

14 days+
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Benefits offered by this job

Competitive compensation
Comprehensive benefits
Flexible work arrangements

Job summary

A cutting-edge AI company in California is looking for a Member of Technical Staff for Kernel/Compiler/Communication. This critical role requires strong expertise in CUDA and GPU optimization, along with 5+ years of experience in performance engineering. The ideal candidate will design high-performance kernels and optimize systems for large GPU clusters, contributing to the next generation of AI solutions. The company offers competitive compensation, equity, and a collaborative work environment.

Qualifications

  • 5+ years of experience in systems, compiler, or performance engineering.
  • Strong expertise in CUDA or accelerator programming.
  • Deep understanding of GPU architecture and memory hierarchy.
  • Experience writing or optimizing high-performance kernels.
  • Strong background in compilers, runtimes, or code generation.
  • Experience with distributed communication libraries (NCCL, MPI, RCCL, etc.)
  • Solid knowledge of networking and interconnect technologies
  • Proficiency in C++ and Python
  • Strong debugging and profiling skills at system level

Responsibilities

  • Design and implement high-performance kernels for AI workloads.
  • Optimize compiler and runtime stacks for ML systems.
  • Improve communication efficiency across large GPU clusters.
  • Reduce latency and increase throughput for distributed workloads.
  • Profile and eliminate system bottlenecks across the stack.
  • Collaborate with training and inference teams on performance optimization
  • Develop tooling for profiling and performance analysis
  • Contribute to long‑term architecture for performance‑critical systems
  • Push the limits of hardware–software co‑design

Skills

CUDA or accelerator programming
GPU architecture and memory hierarchy
C++
Python
Debugging and profiling skills at system level
Debugging / profiling
Distributed communication libraries

Tools

NCCL
MPI
RCCL
TVM
Triton
MLIR
NVLink
InfiniBand
RDMA

Job description

Member of Technical Staff — Kernel / Compiler / Communication
About the Role

RadixArk is seeking aMember of Technical Staff — Kernel / Compiler / Communicationto push the limits of performance for frontier AI systems.

You will work at the lowest layers of the stack — kernels, runtimes, compilers, and communication libraries — to unlock maximum efficiency from modern accelerators and interconnects.

This role is critical to scaling training and inference across thousands of GPUs, where microseconds and memory bandwidth matter. Your work will directly shape the performance envelope of next-generation AI systems.

This is a deeply technical role for engineers who enjoy working close to hardware and solving performance problems that most engineers never encounter.

Requirements

5+ years of experience in systems, compiler, or performance engineering

Strong expertise in CUDA or accelerator programming

Deep understanding of GPU architecture and memory hierarchy

Experience writing or optimizing high-performance kernels

Strong background in compilers, runtimes, or code generation

Experience with distributed communication libraries (NCCL, MPI, RCCL, etc.)

Solid knowledge of networking and interconnect technologies

Proficiency in C++ and Python

Strong debugging and profiling skills at system level

Strong Plus

Experience with Triton, TVM, XLA, or MLIR

Experience building compiler passes or IR transformations

Familiarity with NVLink, InfiniBand, or RDMA

Experience optimizing collective communication at scale

Background in HPC or performance‑critical systems

Contributions to kernel/compiler/ML systems open source

Experience scaling workloads to 1000+ GPUs

Experience with mixed‑precision or quantized kernels

Responsibilities

Design and implement high-performance kernels for AI workloads

Optimize compiler and runtime stacks for ML systems

Improve communication efficiency across large GPU clusters

Reduce latency and increase throughput for distributed workloads

Profile and eliminate system bottlenecks across the stack

Collaborate with training and inference teams on performance optimization

Develop tooling for profiling and performance analysis

Contribute to long‑term architecture for performance‑critical systems

Push the limits of hardware–software co‑design

About RadixArk

RadixArk is an infrastructure‑first AI company built by engineers who have shipped production AI systems, created SGLang (20K+ GitHub stars, the fastest open LLM serving engine), and developed Miles, our large‑scale RL framework.

We build world‑class systems for AI training and inference and collaborate with frontier AI labs and cloud providers.

Our team has optimized kernels serving billions of tokens daily and designed distributed systems coordinating 10,000+ GPUs.

Join us to build the performance foundation of next‑generation AI.

Compensation

We offer competitive compensation with meaningful equity, comprehensive benefits, and flexible work arrangements. Compensation depends on location, experience, and level.

RadixArk is an Equal Opportunity Employer and welcomes candidates from all backgrounds.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Member of Technical Staff — Heterogenous Hardware
Member of Technical Staff — Heterogenous Hardware

RadixArk • Palo Alto (CA)

On-site
USD 180,000 - 240,000
Competitive compensation
Equity
Flexible work arrangements
+1
Member of Technical Staff — Inference
Member of Technical Staff — Inference

Dormont Manufacturing Co • Palo Alto (CA)

On-site
USD 190,000 - 260,000
Competitive compensation
Meaningful equity
Comprehensive benefits
+1
Member of Technical Staff — Performance
Member of Technical Staff — Performance

RadixArk • Palo Alto (CA)

On-site
USD 120,000 - 160,000
Competitive compensation
Comprehensive health benefits
Flexible work arrangements
Member of Technical Staff — Cluster / Platform
Member of Technical Staff — Cluster / Platform

RadixArk • Palo Alto (CA)

On-site
USD 200,000 - 400,000
Member of Technical Staff — Developer Technology
Member of Technical Staff — Developer Technology

RadixArk • Palo Alto (CA)

On-site
USD 180,000 - 240,000
Equity
Flexible work arrangements
Competitive benefits
Member of Technical Staff — Training
Member of Technical Staff — Training

RadixArk • Palo Alto (CA)

On-site
USD 120,000 - 160,000
Comprehensive benefits
Flexible work arrangements
Member of Technical Staff — Reliability-CI Infrastructure
Member of Technical Staff — Reliability-CI Infrastructure

RadixArk • Palo Alto (CA)

On-site
USD 120,000 - 150,000
Equity opportunities
Comprehensive health benefits
Flexible work arrangements
Member of Technical Staff — Inference-TPU
Member of Technical Staff — Inference-TPU

RadixArk • Palo Alto (CA)

On-site
USD 180,000 - 250,000
Significant founding team equity
Comprehensive health benefits
Flexible work arrangements
Member of Technical Staff — Backend/API Platform
Member of Technical Staff — Backend/API Platform

RadixArk • Palo Alto (CA)

On-site
USD 120,000 - 160,000
Equity
Comprehensive health benefits
Flexible work arrangements
Member of Technical Staff — Developer Experience
Member of Technical Staff — Developer Experience

RadixArk • Palo Alto (CA)

On-site
USD 150,000 - 230,000
Equity
Benefits package