Inference Systems Engineer - High-Performance AI

SpaceXAI

Palo Alto (CA)

On-site

USD 180,000 - 440,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Equity
Medical coverage
Vision coverage
Dental coverage
401(k) retirement plan

Job summary

SpaceXAI is building a high-performance inference platform used by Grok at scale. As a Member of Technical Staff - Inference, you will design and optimize distributed infrastructure for model serving, including global KV caching, batching, and auto-scaling, with deep GPU kernel work and code generation.

You will own everything from distributed infrastructure to low-level optimizations, shaping how fast and reliably users interact with Grok across millions of requests per day.

Qualifications

  • BASIC QUALIFICATIONS include deep low-level systems programming in C/C++ or Rust.
  • Experience with large-scale, high-concurrent production serving.
  • Experience with GPU inference engines (e.g., vLLM, SGLang, Triton, TensorRT-LLM).
  • Strong background in system optimizations: batching, caching, load balancing, parallelism.
  • Algorithmic inference optimizations: quantization, speculative decoding, distillation, low-precision numerics.
  • Experience with testing, benchmarking, and reliability of inference services.
  • Experience designing and implementing CI/CD infrastructure for inference.

Responsibilities

  • Architect and implement scalable distributed infrastructure for model serving (load balancing, auto-scaling, batch scheduling, global KV cache).
  • Optimize latency and throughput of model inference under real production workloads.
  • Build reliable, high-concurrency serving systems that serve billions of users with high uptime, low error rate, and excellent tail latency.
  • Benchmark, fine-tune, and accelerate inference engines (including low-level GPU kernel work and code generation).
  • Develop custom tools to trace, replay, and fix issues across the full stack — from orchestration down to GPU kernels.
  • Create robust CI/CD infrastructure for seamless endpoint deployment, image publishing, and inference engine updates.
  • Accelerate research on scaling test-time compute, RL rollout, and model-hardware co-design for next-generation systems.

Skills

C/C++/Rust
Large-scale serving
GPU inference engines
System optimizations
Inference optimizations
Testing & benchmarking
CI/CD for inference

Job description

SpaceXAI is building a high-performance inference platform used by Grok at scale. As a Member of Technical Staff - Inference, you will design and optimize distributed infrastructure for model serving, including global KV caching, batching, and auto-scaling, with deep GPU kernel work and code generation.

You will own everything from distributed infrastructure to low-level optimizations, shaping how fast and reliably users interact with Grok across millions of requests per day.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Staff Engineer - Large-Scale Model Inference & Systems
Staff Engineer - Large-Scale Model Inference & Systems

Xai • Palo Alto (CA)

On-site
USD 180,000 - 440,000
Staff Engineer – High-Performance Model Inference
Staff Engineer – High-Performance Model Inference

Pantera Capital • Palo Alto (CA)

On-site
USD 180,000 - 440,000
Equity
Medical coverage
Vision coverage
+5
Software Engineer - Training/Inference (C++)
Software Engineer - Training/Inference (C++)

SpaceXAI • Palo Alto (CA)

On-site
USD 180,000 - 440,000
Equity
Medical coverage
Vision coverage
+2
Software Engineer - Training/Inference (C++)
Software Engineer - Training/Inference (C++)

Xai • Palo Alto (CA)

On-site
USD 180,000 - 440,000
Inference Systems Engineer, High-Throughput AI Serving
Inference Systems Engineer, High-Throughput AI Serving

Future Ventures • Palo Alto (CA)

On-site
USD 135,000 - 160,000
Comprehensive medical, vision, dental coverage
401(k) retirement plan
Paid parental leave
+1
Software Engineer - Training/Inference (C++)
Software Engineer - Training/Inference (C++)

Pantera Capital • Palo Alto (CA)

On-site
USD 180,000 - 440,000
Equity
Medical coverage
Vision coverage
+5
Platform Infra Engineer (Rust/C++, Kubernetes)
Platform Infra Engineer (Rust/C++, Kubernetes)

SpaceXAI • Bellevue (WA)

On-site
USD 180,000 - 440,000
Senior AI Infra Engineer: High-Performance Inference
Senior AI Infra Engineer: High-Performance Inference

Ddn • Sacramento (CA)

On-site
USD 140,000 - 200,000
Platform Infra Engineer (Rust/C++/K8s)
Platform Infra Engineer (Rust/C++/K8s)

Pantera Capital • Palo Alto (CA)

On-site
USD 180,000 - 440,000
equity
comprehensive medical
vision
+5
Platform Infra Engineer (Rust/C++) — Distributed Systems
Platform Infra Engineer (Rust/C++) — Distributed Systems

SpaceXAI • Palo Alto (CA)

On-site
USD 180,000 - 440,000
Equity
Medical coverage
401(k) plan
+1