Role overview
Build and operate the reliability, delivery, automation, and security foundations that allow product teams to release software quickly and safely. Working within a productivity engineering and self-service platform team, this role owns production reliability practices, incident response, observability, infrastructure automation, Kubernetes operations, deployment safety, and security controls embedded in CI/CD.
Responsibilities
- Define and evolve SLI, SLO, and error-budget practices, using reliability data to influence priorities and product decisions.
- Lead incident response, facilitate post-incident reviews, and convert findings into durable system improvements.
- Build and maintain observability across metrics, logs, and traces while improving signal quality and reducing alert fatigue.
- Design and operate resilient infrastructure with Infrastructure as Code, including capacity planning and cloud-cost optimization.
- Manage production Kubernetes and container workloads and support safe deployment methods such as canary releases, progressive rollouts, and rapid rollback.
- Integrate and tune SAST, DAST, SCA, dependency scanning, and related security controls in delivery pipelines.
- Implement policy-as-code to prevent unsafe infrastructure and Kubernetes changes at admission time.
- Maintain vulnerability triage and remediation service levels, improve on-call sustainability, and coach engineers on operational practices.
Requirements
- 5+ years in site reliability, platform, or infrastructure engineering with senior ownership of production systems.
- Strong programming ability in Go, Python, TypeScript, or a similar language for automation and production tooling.
- Hands-on experience with a major cloud platform, Kubernetes, and Infrastructure as Code; AWS and Terraform experience are useful.
- Proven experience leading incident response and implementing SLO-driven reliability practices.
- Working knowledge of observability tooling; experience with Datadog is useful.
- Practical experience securing CI/CD through scanning, dependency controls, or policy-as-code.
- Understanding of cloud security fundamentals, including IAM, least privilege, guardrails, and secrets management.
- Strong judgment and communication skills when raising reliability or security issues across engineering teams.
Nice to have
- Experience with policy-as-code frameworks such as OPA/Rego, Kyverno, or Conftest.
Benefits and work setup
The source describes comprehensive healthcare, retirement savings with employer matching, paid family leave, fertility support, mental health resources, wellness and technology allowances, performance-related rewards, learning opportunities, mentorship, and an inclusive workplace culture.