Get more replies from employers
Send a job-specific resume in minutes.
Neocloud seeks a founding member to build and operate GPU-focused compute infrastructure in San Francisco. You will directly influence the platform that powers AI workloads, with ownership and visibility from the founders.
This onsite role relocates to San Francisco as needed. You will stand up production Kubernetes/Slurm clusters with custom GPU orchestration, design storage and networking for high-scale workloads, and own incident response end-to-end.
Join a Stealth Neocloud building GPU compute infrastructure from the ground up in San Francisco. As one of the founding team, the systems you design and build this year are the systems the company runs on. You'll have direct access to and collaboration with the founders, high ownership, high visibility, and no bureaucracy standing between you and the rack.
Responsibilities:
Stand up and operate production Kubernetes and/or Slurm clusters with custom GPU orchestration (scheduling logic, topology-aware placement, GPU lifecycle automation) built or substantially customised by you, not tooling run out of the box.
Design and implement the storage architecture (object storage, NVMe clusters, high-bandwidth networking) that the platform depends on, tuning it for real customer workloads at scale.
Work directly with customers to translate workload requirements into infrastructure decisions: sizing clusters, configuring storage and networking, and adjusting orchestration to fit what they actually need to run.
Own cluster reliability and production incident response end-to-end, documenting runbooks and operational processes so the next hire doesn't start from zero.
Skills/Must have:
Operating Kubernetes and/or Slurm clusters at scale with custom GPU orchestration, scheduling logic, and lifecycle automation built or substantially customised in-house.
Deep hands-on experience with distributed object storage, NVMe storage clusters, high-bandwidth networking fabrics, and storage architectures optimised for large-scale AI workloads.
Designing and operating scalable inference platforms for deploying, serving, and optimising large AI models with low-latency, high-throughput GPU acceleration.
End-to-end ownership of distributed system reliability, incident response, and operational maturity; comfort being the first call when production breaks.
Based in or willing to relocate to San Francisco (onsite).
Founding-level ownership and visibility.
Direct access to and collaboration with the founders.
On-site role in San Francisco; relocation provided if needed.