SRE Manager: Cloud Reliability & Observability

Twist Bioscience Corporation

San Francisco (CA)

Hybrid

USD 210,000 - 243,000

Full time

6 days ago
Be an early applicant
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Job summary

Twist Bioscience Corporation seeks an experienced Manager, Site Reliability Engineering to lead Site Operations and infrastructure initiatives in a hybrid role based in South San Francisco. You will drive reliability, scalability, and security across cloud platforms and critical applications.

You will collaborate with engineering teams to improve deployment processes, implement CI/CD, and foster a culture of reliability while mentoring SRE engineers.

Qualifications

  • 5–7+ years of experience in Site Operations, Site Reliability Engineering, Infrastructure Engineering, or related operational roles.
  • Prior experience leading or mentoring engineering or operations teams.
  • Strong hands‑on experience with cloud infrastructure platforms (AWS, Azure, and/or GCP).
  • Experience managing Kubernetes (K8s) environments in production.
  • Strong understanding of networking concepts including DNS, load balancing, firewalls, routing, and CDN technologies.
  • Experience with Akamai CDN administration and optimization.
  • Experience with observability and monitoring platforms, including Splunk.
  • Experience with storage infrastructure and cloud-native storage solutions.
  • Strong scripting and automation experience (Python, Bash, Terraform, or similar).
  • Experience implementing and maintaining CI/CD pipelines.
  • Strong troubleshooting and incident management skills in complex distributed systems.
  • Excellent communication and cross‑functional collaboration skills.

Responsibilities

  • Lead and manage the Site Operations / SRE function.
  • Own cloud infrastructure architecture, operations, scalability, and optimization.
  • Ensure high availability and reliability of production applications and services.
  • Drive operational excellence through automation, monitoring, and incident management.
  • Develop and maintain observability platforms including logging, metrics, alerting, and tracing.
  • Manage Kubernetes-based infrastructure and container orchestration environments.
  • Oversee networking, DNS, CDN, and storage infrastructure across cloud environments.
  • Collaborate with software engineering teams to improve deployment processes, system resilience, and performance.
  • Establish and improve CI/CD pipelines and infrastructure-as-code practices.
  • Lead incident response, root cause analysis, and post-incident remediation efforts.
  • Manage vendor and platform relationships where applicable.
  • Mentor and develop SRE / Site Operations engineers and foster a culture of reliability and continuous improvement.
  • Define and implement infrastructure security and compliance best practices.

Skills

Leadership
Cross-functional collaboration
Incident management
Cloud infrastructure
Observability

Tools

Kubernetes
Akamai CDN
Splunk
Terraform
Python
Bash

Job description

Twist Bioscience Corporation seeks an experienced Manager, Site Reliability Engineering to lead Site Operations and infrastructure initiatives in a hybrid role based in South San Francisco. You will drive reliability, scalability, and security across cloud platforms and critical applications.

You will collaborate with engineering teams to improve deployment processes, implement CI/CD, and foster a culture of reliability while mentoring SRE engineers.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

SRE Manager - Cloud Reliability Leader (Hybrid)
SRE Manager - Cloud Reliability Leader (Hybrid)

Twist Bioscience Corporation • South San Francisco (CA)

Hybrid
USD 210,000 - 243,000
SRE Manager: Lead Reliability & Observability at Scale
SRE Manager: Lead Reliability & Observability at Scale

Iac/interactivecorp • Sacramento (CA)

On-site
USD 150,000 - 210,000
Collaborative work environment
Commitment to carbon‑reduction mission
Flex schedule
+2
SRE Engineering Manager — Lead Reliability, Remote Flexible
SRE Engineering Manager — Lead Reliability, Remote Flexible

DevOpsChat • California (MO)

Hybrid
USD 140,000 - 220,000
Health insurance
Professional development opportunities
Senior Production SRE: Cloud & On-Prem Reliability
Senior Production SRE: Cloud & On-Prem Reliability

Weights & Biases • New York (NY)

On-site
USD 140,000 - 180,000
Medical Insurance
Dental Insurance
Vision Insurance
+15
Senior SRE Leader: Cloud, Reliability & Scale
Senior SRE Leader: Cloud, Reliability & Scale

AVG • Tempe (AZ), Northern (KY)

Hybrid
USD 140,000 - 190,000
Remote SRE Manager: Lead Reliability & Automation
Remote SRE Manager: Lead Reliability & Automation

NationsBenefits, LLC • Plantation (FL)

Remote
USD 140,000 - 180,000
Fully remote
Unlimited PTO
Competitive compensation
+1
Staff SRE – Cloud-Native Reliability & SecOps
Staff SRE – Cloud-Native Reliability & SecOps

Crunchyroll • San Francisco (CA)

On-site
USD 233,000 - 292,000
Performance bonus
Flexible time off
Medical insurance
+5
Remote SRE Manager: Cloud Reliability & DevOps Lead
Remote SRE Manager: Cloud Reliability & DevOps Lead

Deepwatch • Tampa (FL)

Hybrid
USD 178,000 - 213,000
Medical, dental, vision insurance
Flexible Time Off
Professional development benefits
+1
Remote SRE Manager: Reliability & Scale for Fintech Health
Remote SRE Manager: Reliability & Scale for Fintech Health

NationsBenefits • Plantation (FL)

On-site
USD 140,000 - 190,000
Unlimited PTO
Fully remote work (US-based)
Senior SRE: Cloud, Kubernetes & Automation
Senior SRE: Cloud, Kubernetes & Automation

Socure • Carson City (NV)

On-site
USD 150,000 - 190,000