ABOUT THE ROLE
You will keep the studio's products reliable, scalable, and observable as they grow across the portfolio. Bridging engineering and operations, you build the infrastructure, automation, and guardrails that let teams ship fast and safely. You treat operations as a software problem — automating toil, codifying infrastructure, and designing systems that stay up under real-world load. You bring calm and rigor to incidents and a bias toward prevention.
KEY RESPONSIBILITIES
- Design, build, and maintain the cloud infrastructure and deployment pipelines that products across the portfolio run on.
- Define and manage infrastructure as code, and automate provisioning, configuration, and releases.
- Establish observability — logging, metrics, tracing, and alerting — so teams can understand and trust their systems in production.
- Set and track reliability targets (SLOs/SLIs) and drive improvements to availability, latency, and performance.
- Lead incident response: detect, triage, mitigate, and run blameless post-mortems that prevent recurrence.
- Design for scalability, resilience, and disaster recovery, including backups, failover, and capacity planning.
- Harden infrastructure and partner with security to enforce access controls, secrets management, and safe defaults.
- Reduce operational toil through automation, and coach product teams on reliability and operational best practices.
REQUIRED QUALIFICATIONS
- 4+ years in SRE, DevOps, platform, or infrastructure engineering roles running production systems.
- Strong hands-on experience with a major cloud provider and with containerization and orchestration (Docker, and Kubernetes or similar).
- Proficiency with infrastructure as code (e.g., Terraform) and CI/CD tooling.
- Solid scripting/programming ability for automation (e.g., Python, Go, or Bash).
- Practical experience with observability stacks (metrics, logging, tracing) and incident management.
- Sound understanding of networking, Linux systems, and secure operational practices.
- Calm, methodical approach to troubleshooting and production incidents.
PREFERRED QUALIFICATIONS
- Experience supporting multiple products or a diverse technology stack at once.
- Familiarity with defining SLOs/SLIs and running an error-budget-driven practice.
- Exposure to database operations, message queues, or workflow orchestration in production.
- Experience with cost optimization and capacity planning across cloud environments.
*These descriptions are intended as adaptable templates. Compensation, location, employment type, and studio-specific details can be added per posting.*