Get more replies from employers
Send a job-specific resume in minutes.
CompanyThunder Compute is building a cloud infrastructure stack for GPU virtualization, scaling from design to production as part of an ambitious, hands-on team in San Francisco.
You will own the Go-based cloud platform, Kubernetes orchestration, and core services for provisioning, scheduling, billing, and secure access across multiple data centers and cloud providers. This is a high-impact role requiring pragmatic engineering and strong collaboration.
CompanyThunder Compute is building the VMware for GPUs. We have raised over $17.5M from Matrix Partners, Y Combinator, and leading angels from Coreweave, Microsoft, Cognition, and Anthropic. Deployed GPU fleets are currently only 5-20% utilized. Leading solutions for underutilization sit at the workload layer and are therefore only able to optimize specific use cases. We believe the ideal cluster optimization solution must be invisible to developers and compatible with all workloads; hence, it must sit at the systems layer. We are a team of systems researchers productionizing cutting-edge GPU virtualization research to build this general-purpose optimization layer. Concretely, our virtualization library abstracts GPUs across TCP networking. We use a userspace shim library, loaded through LD_PRELOAD, to intercept CUDA calls and send them over gRPC to a host server connected to a physical GPU elsewhere in the data center. This enables something like "Ceph for GPUs": GPUs become network resources that can be abstracted, pooled, and dynamically allocated across a cluster to improve utilization without requiring developers to modify their workloads.
Your work will focus on building the cloud infrastructure surrounding our GPU virtualization layer. This includes the Go backbone of our cloud platform, Kubernetes-based orchestration, production reliability, networking, storage, billing infrastructure, and the systems used to deploy and operate GPU capacity at scale. You will take ownership of complex infrastructure from early design through production deployment. Example projects may include:
You will spend your days bouncing between the weeds of complicated production infrastructure that is live and used by customers. One week, you may be debugging a networking failure across a Kubernetes cluster; the next, you may be redesigning the provisioning system to make deployments faster and more reliable. This work is not easy. It blends the hardest parts of cloud infrastructure, distributed systems, and production engineering. We look for exceptional engineering talent, strong work ethic, and extreme attention to detail. We must move quickly while shipping high-quality, reliable infrastructure.
You will join early enough to meaningfully shape the architecture, engineering standards, and technical direction of the company. You will work directly with the founders on a category-defining systems problem, with a short path between writing code and seeing it run in production. The infrastructure you build will operate a new foundational layer for GPU computing. Unlike at a large company, you will not be restricted to one small component of a much larger system. You will own broad, technically difficult areas of the platform and have the opportunity to grow into senior technical and engineering leadership as the company scales.
You will report to co-founder and CTO Brian Model, formerly a Quantitative Developer at Citadel Securities.
This role is full-time and in person, five days per week, at our office in downtown San Francisco.
Relocation support and visa sponsorship are available.