Turn this role into an interview — a resume and cover letter built around what this employer wants.
Weights & Biases is seeking a seasoned SRE/DevOps engineer to design and operate a reliability platform across multi-cloud and on‑prem environments, enabling safe and fast software delivery.
You will automate toil, enhance observability with Prometheus, Grafana, and Datadog, and lead incident response across cloud providers to minimize on‑call burden.
The role emphasizes systems programming, IaC leadership, and collaborating with product teams to ship resilient infrastructure at scale.
Design and operate a reliability platform across multi-cloud and on-premises environments to enable safe and fast software delivery. Focus on automating toil, improving observability, and managing incident response systems to reduce on-call burden.
Requires extensive experience in production infrastructure, expertise in major public clouds, and proficiency in Kubernetes and Infrastructure as Code. Candidates must be skilled in systems programming and have a proven track record of technical leadership in SRE/DevOps roles.