Flexible work arrangement (remote or San Francisco office)
Full visa sponsorship and relocation support
Professional development budget
Regular team off-sites and conference attendance
Opportunity to shape decentralized AI and RL
Job summary
Prime Intellect is looking for a skilled ML Systems Engineer to build and optimize LLM serving infrastructure and inference systems. This hybrid role involves contributing to the scalability of their reinforcement learning training. Successful candidates will have over 3 years of experience in building ML services, strong knowledge of Python and cloud platforms, and a desire to work on cutting-edge AI infrastructure. They offer a cash compensation range of $150-300k with equity and full relocation support.
Qualifications
3+ years building and running large-scale ML/LLM services with clear latency/availability SLOs.
Hands-on with vLLM, SGLang, TensorRT‑LLM.
Familiarity with distributed and disaggregated serving infrastructure such as NVIDIA Dynamo.
Deep understanding of prefill vs. decode, KV-cache behavior, batching, and speculative decoding.
Comfortable debugging CUDA/NCCL, drivers/kernels, and storage.
Responsibilities
Build infrastructure to serve LLMs efficiently at scale.
Optimize and integrate inference systems into our RL training stack.
Design placement and scheduling algorithms for heterogeneous accelerators.
Implement multi-region/zone failover and traffic shifting.
Profile kernels, memory bandwidth and transport; apply quantization techniques.
Skills
Building ML Systems at Scale
Inference Backends
Distributed Serving Infra
Inference Internals
Full-Stack Debugging
Python
PyTorch
Cloud & Automation
Kubernetes
GPU & Networking
Job description
Prime Intellect is looking for a skilled ML Systems Engineer to build and optimize LLM serving infrastructure and inference systems. This hybrid role involves contributing to the scalability of their reinforcement learning training. Successful candidates will have over 3 years of experience in building ML services, strong knowledge of Python and cloud platforms, and a desire to work on cutting-edge AI infrastructure. They offer a cash compensation range of $150-300k with equity and full relocation support.