Turn this role into an interview — a resume and cover letter built around what this employer wants.
CoreWeave is seeking a Staff Software Engineer to join the Compute Architecture group. You will help build software systems that operate the backbone of large-scale GPU data centers, focusing on Go-based distributed services and infrastructure automation.
You will design, build, and operate services that manage lifecycle, health, and firmware state across fleets of GPU servers, with emphasis on reliability, scalability, and observability. Strong leadership and mentoring are expected.
Skilled in applying a data-driven approach to reliability, optimization, and continuous improvementB.S., M.S., or PhD in Computer Science or related field, or equivalent experience8+ years of software engineering experience with a strong focus on infrastructure, cloud engineering, and distributed databases—particularly within large-scale datacenter and cloud environmentsExpertise in Go and proven experience building REST/gRPC APIs for mission-critical platformsExcellent communicator able to work effectively with both technical and non-technical stakeholdersTrack record of leading incident response, postmortems, and driving robust service reliabilityProven success in mentoring engineers, leading technical projects, and influencing engineering strategy across teamsHands-on experience with observability stacks (Prometheus, Grafana, PromQL), CI/CD pipelines, and operating large fleets of GPU serversStrong background in architecting and scaling cloud-native Kubernetes infrastructure and distributed servicesExperience contributing to and collaborating with open source communitiesWorking knowledge of Kafka, ClickHouse and CRDBDMTF, RedFish APIs, and GPU servers