Get more replies from employers
Send a job-specific resume in minutes.
CompanyThunder Compute is building a general-purpose GPU virtualization layer, enabling GPUs to be pooled and allocated like networked resources. You will own core C++ systems, focusing on latency, reliability, and production readiness, with research into new allocation and oversubscription techniques.
You will work on profiling, debugging, and extending support across CUDA apps and hardware, collaborating with customers and cross-functional teams to deliver scalable GPU infrastructure in
CompanyThunder Compute is building the VMware for GPUs. We have raised over $17.5M from Matrix Partners, Y Combinator, and leading angels from Coreweave, Microsoft, Cognition, and Anthropic. Deployed GPU fleets are currently only 5-20% utilized. Leading solutions for underutilization sit at the workload layer and are therefore only able to optimize specific use cases. We believe the ideal cluster optimization solution must be invisible to developers and compatible with all workloads; hence, it must sit at the systems layer. We are a team of systems researchers productionizing cutting-edge GPU virtualization research to build this general-purpose optimization layer.
Concretely, our virtualization library abstracts GPUs across TCP networking. We use a userspace shim library, loaded through LD_PRELOAD, to intercept CUDA calls and send them over gRPC to a host server connected to a physical GPU elsewhere in the data center. This enables something like "Ceph for GPUs": GPUs become network resources that can be abstracted, pooled, and dynamically allocated across a cluster to improve utilization without requiring developers to modify their workloads.
Your work will focus on building the core C++ systems behind our virtualization layer. This includes low-latency performance optimization, distributed systems debugging, production reliability, and research into new techniques for improving GPU utilization.
You will take ownership of complex systems from early experimentation through production deployment. Example projects may include:
You will join early enough to meaningfully shape the architecture, engineering standards, and technical direction of the company. You will work directly with the founders on a category-defining systems problem, with a short path between writing code and seeing it run in production. The systems you build will form the foundation of a new infrastructure layer for GPU computing.
You will report to co-founder and CTO Brian Model, formerly a Quantitative Developer at Citadel Securities. This role is full-time and in person, five days per week, at our office in downtown San Francisco. Relocation support and visa sponsorship are available.