Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.
Coreweave is seeking a Senior Operations Engineer to join the MetalDev Operations team. You will own monitoring, incident response, and tooling, focusing on reliability and automation in a cutting-edge data center environment.
You will work with hardware and firmware vendors, collaborate with hardware engineering and fleet operations, and drive improvements in runbooks and on-call processes to reduce toil and incidents.
Experience deploying and supporting containerized applications in Kubernetes environmentsExcellent written and verbal communication skills, particularly during high-impact incidents5+ of experience in cloud operations, site reliability engineering (SRE), infrastructure operations, or a related technical fieldExperience troubleshooting complex issues across software services, operating systems, networks, and physical infrastructureStrong documentation skills and attention to detailExperience participating in an on‑call rotation supporting production servicesExperience using Prometheus, Grafana, and PromQL for monitoring, alerting, and troubleshootingStrong knowledge of Linux system administration and internals and scriptingWorking knowledge of Kubernetes and at least one public cloud platform, such as AWS or GCPStrong analytical and problem‑solving skills, with a methodical approach to troubleshootingExperience with incident management practices, including incident response, escalation, and post‑incident review processesYou like collaborating across hardware, platform, and product teams to solve complex, ambiguous problemsWe believe in investing in our people, and value candidates who can bring their own diversified experiences to our teams – even if you aren’t a 100% skill or experience match. Here are a few qualities we’ve found compatible with our team. If some of this describes you, we’d love to talkYou enjoy working close to the hardware and are curious about how GPUs, servers, and data centers fit togetherYou thrive in infrastructure environments where reliability, performance, and automation matter as much as featuresExperience with server hardware, BMCs, Redfish, IPMI, or hardware‑management servicesUnderstanding of Python or GolangExperience working in data center environments, including server racks, power‑distribution equipment, and cooling systemsExperience troubleshooting server provisioning, reboot, provisioning, or lifecycle‑management failuresExperience collaborating directly with hardware or firmware vendors to qualify and validate fixesBachelor’s degree in computer science, engineering, or a related discipline—or equivalent practical experienceFamiliarity with high‑performance computing, GPU infrastructure, DPUs, or large‑scale AI clusters