Get more replies from employers
Send a job-specific resume in minutes.
SpaceXAI is building a high-performance inference platform serving Grok at global scale. You will design and optimize distributed model-serving systems, owning from global KV caches to low-level GPU kernel work and code generation.
This role drives latency, throughput, and reliability for billions of users, with opportunities to push innovations in batching, quantization, and speculative decoding on next-gen hardware.
SpaceXAI is building a high-performance inference platform serving Grok at global scale. You will design and optimize distributed model-serving systems, owning from global KV caches to low-level GPU kernel work and code generation.
This role drives latency, throughput, and reliability for billions of users, with opportunities to push innovations in batching, quantization, and speculative decoding on next-gen hardware.