Turn this role into an interview — a resume and cover letter built around what this employer wants.
Runpod, a remote-first AI developer cloud, seeks a Site Reliability Engineer to own reliability, observability, and operations for a distributed platform. You will design SLIs/SLOs, drive incident response, and automate deployments while strengthening production readiness.
You’ll work with cross-functional teams to reduce toil, improve MTTR, and ensure scalable performance in a fast-growing environment. Strong scripting and GPU awareness are valued.
Runpod, a remote-first AI developer cloud, seeks a Site Reliability Engineer to own reliability, observability, and operations for a distributed platform. You will design SLIs/SLOs, drive incident response, and automate deployments while strengthening production readiness.
You’ll work with cross-functional teams to reduce toil, improve MTTR, and ensure scalable performance in a fast-growing environment. Strong scripting and GPU awareness are valued.