A technology recruitment agency seeks a professional to lead the architecture of a distributed inference platform within a stealth-mode hyperscale data center startup. Applicants must have over 5 years of experience with large-scale distributed systems and strong proficiency in languages such as Python, Go, or Rust, along with a solid understanding of GPU software stacks. This unique opportunity offers an annual salary of $300,000 and includes equity benefits.
Qualifications
5+ years’ experience building large-scale, fault-tolerant distributed systems.
Strong understanding of GPU software stacks (CUDA, Triton, NCCL) and Kubernetes.
Practical experience with model-serving frameworks such as vLLM, SGLang, TensorRT-LLM.
Responsibilities
Take ownership of the inference platform architecture.
Design, build, and optimise distributed inference systems.
Integrate, tune, and operate inference engines across multiple model types.
Develop APIs and orchestration layers for multi-tenant and dedicated deployments.
Skills
Building large-scale distributed systems
Proficiency in Python, Go, Rust
Understanding of GPU software stacks
Experience with model-serving frameworks
Knowledge of performance optimisation techniques
Familiarity with Infrastructure-as-Code tools
Job description
A technology recruitment agency seeks a professional to lead the architecture of a distributed inference platform within a stealth-mode hyperscale data center startup. Applicants must have over 5 years of experience with large-scale distributed systems and strong proficiency in languages such as Python, Go, or Rust, along with a solid understanding of GPU software stacks. This unique opportunity offers an annual salary of $300,000 and includes equity benefits.