AI Inference Systems Engineer (High-Throughput, Low-Latency)

SpaceX

Palo Alto (CA)

On-site

USD 135,000 - 210,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Stock options
Excellent medical coverage
401(k) plan

Job summary

SpaceX in Palo Alto is seeking a Software Engineer, Inference (AI Data Engineering) to design and optimize large-scale model serving systems. You will own distributed infrastructure, drive high-throughput, low-latency inference for SpaceX's critical workloads.

You will work on GPU kernels, quantization, and acceleration techniques, collaborating with diverse teams to deliver reliable, scalable AI infrastructure that powers our ambitious missions. On-site work in Palo Alto is required.

Qualifications

  • Bachelor's degree in computer science or related field, or 2+ years of professional software experience.
  • Experience designing, implementing, and maintaining reliable, horizontally scalable distributed systems.
  • 1+ years of full stack or backend production systems experience.
  • 1+ years of experience with Rust or C++.

Responsibilities

  • Develop highly reliable, high-throughput inference systems serving internal AI models.
  • Architect scalable distributed infrastructure for model serving with autoscaling and batching.
  • Optimize latency and throughput under production workloads using GPU techniques.
  • Build reliable, high-concurrency serving systems with observability and uptime.
  • Own components like request routing, SDK development, rate limiting, and scaling.
  • Benchmark and accelerate inference engines (e.g., SGLang, vLLM, TensorRT-LLM).
  • Develop tools for tracing, replaying, and debugging across the stack.

Skills

Rust
C++
Python
Go
Docker
Kubernetes
gRPC
GPU kernels

Education

Bachelor's degree in CS or related
2+ years of professional software experience

Tools

PostgreSQL
ClickHouse
MongoDB

Job description

SpaceX in Palo Alto is seeking a Software Engineer, Inference (AI Data Engineering) to design and optimize large-scale model serving systems. You will own distributed infrastructure, drive high-throughput, low-latency inference for SpaceX's critical workloads.

You will work on GPU kernels, quantization, and acceleration techniques, collaborating with diverse teams to deliver reliable, scalable AI infrastructure that powers our ambitious missions. On-site work in Palo Alto is required.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

AI Inference Engineer - Scalable, Low-Latency Systems
AI Inference Engineer - Scalable, Low-Latency Systems

SPACE EXPLORATION TECHNOLOGIES CORP • Palo Alto (CA), Northern (KY)

Hybrid
USD 135,000 - 210,000
401(k)
Medical, vision and dental coverage
Paid parental leave
+4
Staff Engineer - Large-Scale Model Inference & Systems
Staff Engineer - Large-Scale Model Inference & Systems

Xai • Palo Alto (CA)

On-site
USD 180,000 - 440,000
Software Engineer, Inference (AI Data Engineering)
Software Engineer, Inference (AI Data Engineering)

SPACE EXPLORATION TECHNOLOGIES CORP • Palo Alto (CA), Northern (KY)

Hybrid
USD 135,000 - 210,000
401(k)
Medical, vision and dental coverage
Paid parental leave
+4
Software Engineer - Training/Inference (C++)
Software Engineer - Training/Inference (C++)

Xai • Palo Alto (CA)

On-site
USD 180,000 - 440,000
Inference Infra Engineer: Scale Low-Latency AI Serving
Inference Infra Engineer: Scale Low-Latency AI Serving

Elorian • Palo Alto (CA)

On-site
USD 200,000 - 400,000
Health, dental, and vision benefits
Unlimited PTO
Parental leave
+1
Software Engineer, Inference (AI Data Engineering)
Software Engineer, Inference (AI Data Engineering)

SpaceX • Palo Alto (CA)

On-site
USD 135,000 - 210,000
Stock options
Excellent medical coverage
401(k) plan
AI Data Operations Engineer
AI Data Operations Engineer

AI Chopping Block • Palo Alto (CA)

On-site
USD 144,000 - 270,000
Equity
Medical coverage
Vision coverage
+4
AI Supercomputer Network Engineer
AI Supercomputer Network Engineer

Pantera Capital • Southaven (MS)

On-site
USD 140,000 - 190,000
Senior AI Inference Systems Engineer | GPU Kernels & Runtime
Senior AI Inference Systems Engineer | GPU Kernels & Runtime

NVIDIA • Santa Clara (CA)

On-site
USD 184,000 - 288,000
Senior AI Platform Engineer
Senior AI Platform Engineer

jobs.frontdoordefense.com - Jobboard • Starbase (TX)

On-site
USD 120,000 - 150,000