A complete application in a minute — tailored resume and cover letter, ready to send.
Cloudjobs in San Jose leads the technical vision and development of a managed Kubernetes platform tailored for AI workloads on bare metal. You will collaborate with infrastructure teams to design GPU-aware orchestration, scalable control planes, and high-performance networking to support demanding AI workloads.
The role requires 10+ years in software engineering with at least 5 years in Kubernetes at scale, with expert Go and Python skills and deep GPU orchestration, distributed systems, and
Lead the technical vision and development of a managed Kubernetes platform purpose-built for AI workloads on bare metal. Collaborate across infrastructure teams to design GPU-aware orchestration, high-performance networking, and scalable control planes.
Requirements: Requires 10+ years of software engineering experience with at least 5 years focused on Kubernetes at scale and expert-level proficiency in Go and Python. Must have deep expertise in GPU orchestration, distributed systems, and high-performance Linux networking.
Key Skills: Kubernetes, Go, Python, GPU Orchestration, Distributed Systems, Linux Networking, Bare Metal Infrastructure, NVIDIA GPU Operator, Cloud Native Engineering, Multi-tenancy, Infrastructure as Code, GitOps, Cilium, InfiniBand, RDMA, Prometheus
Benefits: Cash compensation, Equity compensation, Health coverage, Dental coverage, Vision coverage, Wellness stipend, Commuter stipend, 401k Plan with 2% company match, Flexible paid time off