Accelerator Systems Engineer

RadixArk

Palo Alto (CA)

On-site

USD 210,000 - 310,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

RadixArk is seeking a Member of Technical Staff: Accelerator Systems to push the limits of frontier AI performance. You will bring up, optimize, and maintain SGLang, Miles, and the RadixArk stack across GPUs, TPUs, CPUs, and emerging accelerators, porting kernels and runtimes while crafting abstractions for a single codebase to stay fast on all platforms.

You will work directly with silicon and partners on pre-release platforms, solving memory hierarchy challenges and architecture tradeoffs

Qualifications

  • 4+ years of experience in systems, performance, or ML infrastructure engineering.
  • Deep expertise in at least one accelerator programming model (CUDA, ROCm/HIP, Pallas/XLA, Triton, or vendor SDK).
  • Strong understanding of accelerator architecture: memory hierarchy and bandwidth tradeoffs.
  • Experience writing or optimizing high-performance kernels for ML workloads.
  • Experience with distributed execution and communication libraries (NCCL, RCCL, MPI, or equivalents).
  • Proficiency in C++ and Python.
  • Strong debugging and profiling skills at the system level.
  • Track record of performance work that shipped into production.

Responsibilities

  • Bring up RadixArk's inference and training systems on new accelerator platforms and drive them to competitive performance.
  • Design hardware abstractions that let a single codebase stay fast across vendors without forking.
  • Port and optimize kernels across programming models and memory architectures.
  • Build cross-platform benchmarking, profiling, and regression detection so performance claims hold up on every target.
  • Debug numerical divergence and correctness gaps between platforms.
  • Work with vendor engineering teams on pre-release hardware, compiler and driver issues, and roadmap feedback.
  • Partner with kernel, runtime, distributed systems, and product engineers to land performance wins end to end.
  • Serve as the internal source of truth on what each platform is actually good at.
  • Contribute hardware-specific optimizations, benchmarks, and portability work back to open-source SGLang and Miles.

Skills

C++
Python
CUDA
Performance profiling
Distributed systems

Tools

NCCL
RCCL
MPI
MLIR
XLA
Triton

Job description

RadixArk is seeking a Member of Technical Staff: Accelerator Systems to push the limits of frontier AI performance. You will bring up, optimize, and maintain SGLang, Miles, and the RadixArk stack across GPUs, TPUs, CPUs, and emerging accelerators, porting kernels and runtimes while crafting abstractions for a single codebase to stay fast on all platforms.

You will work directly with silicon and partners on pre-release platforms, solving memory hierarchy challenges and architecture tradeoffs

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Staff Accelerator Systems Engineer
Staff Accelerator Systems Engineer

RadixArk • Palo Alto (CA)

On-site
USD 180,000 - 240,000
Competitive compensation
Equity
Flexible work arrangements
+1
Member of Technical Staff — Heterogenous Hardware
Member of Technical Staff — Heterogenous Hardware

RadixArk • Palo Alto (CA)

On-site
USD 180,000 - 240,000
Competitive compensation
Equity
Flexible work arrangements
+1
Member of Technical Staff — Inference-Multi-HardwarePalo Alto, CA
Member of Technical Staff — Inference-Multi-HardwarePalo Alto, CA

RadixArk • Palo Alto (CA)

On-site
USD 210,000 - 310,000
Staff DevTech: Accelerate AI Inference on GPUs
Staff DevTech: Accelerate AI Inference on GPUs

RadixArk • Palo Alto (CA)

On-site
USD 180,000 - 240,000
Member of Technical Staff — Inference-Kernel, Compiler & Communication
Member of Technical Staff — Inference-Kernel, Compiler & Communication

RadixArk • Palo Alto (CA)

On-site
USD 210,000 - 290,000
Competitive compensation
Comprehensive benefits
Flexible work arrangements
Staff, Developer Technology - GPU AI Inference & Training
Staff, Developer Technology - GPU AI Inference & Training

RadixArk • Palo Alto (CA)

On-site
USD 180,000 - 240,000
Equity
Flexible work arrangements
Competitive benefits
Flexible Developer Experience Advocate & Community Builder
Flexible Developer Experience Advocate & Community Builder

RadixArk • Palo Alto (CA)

On-site
USD 150,000 - 230,000
Equity
Benefits package
Member of Technical Staff — Developer TechnologyPalo Alto, CA
Member of Technical Staff — Developer TechnologyPalo Alto, CA

RadixArk • Palo Alto (CA)

On-site
USD 180,000 - 240,000
Member of Technical Staff — Developer Technology
Member of Technical Staff — Developer Technology

RadixArk • Palo Alto (CA)

On-site
USD 180,000 - 240,000
Equity
Flexible work arrangements
Competitive benefits
Member of Technical Staff — Performance
Member of Technical Staff — Performance

RadixArk • Palo Alto (CA)

On-site
USD 120,000 - 160,000
Competitive compensation
Comprehensive health benefits
Flexible work arrangements