Staff SRE & Platform Reliability Architect

Grailbio

Edison (CA)

On-site

USD 169,000 - 224,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Flexible time-off
401(k) with employer match
Medical, dental, vision coverage

Job summary

GRAIL is seeking a Staff Site Reliability / DevOps Engineer to lead the reliability, scalability, and security of our cloud-native platform. You will define infrastructure standards, drive reliability decisions, and mentor teams while building scalable systems.

Onsite at Menlo Park, CA, with planned Fall 2026 move to Sunnyvale, CA. You will own Kubernetes reliability, establish SLO/SLI, and lead incident response across AWS/GCP/Azure environments.

Qualifications

  • BS in CS/Engineering or equivalent experience.
  • 8+ years in SRE/DevOps/platform engineering.
  • Hands-on with major cloud platforms (AWS/GCP/Azure).
  • Experience IaC solutions (Terraform, CloudFormation).
  • Experience designing/operating CI/CD pipelines.
  • Hands-on Kubernetes in production environments.
  • Automation scripting in Python/Go/Bash/PowerShell.
  • Observability and monitoring tooling expertise.
  • Strong networking, security, and distributed systems fundamentals.
  • Experience in regulated environments (ISO 27001/NIST/SOC2/HIPAA).

Responsibilities

  • Design, build, and operate highly available cloud infrastructure across cloud platforms.
  • Architect and maintain scalable CI/CD pipelines and deployment frameworks.
  • Lead infrastructure-as-code adoption and maturity across teams.
  • Own Kubernetes reliability across multi-cluster environments.
  • Establish observability platforms and define SLO/SLI frameworks.
  • Lead incident response, perform RCAs, and implement improvements.
  • Optimize infrastructure for cost, performance, and scalability.
  • Define and enforce DevOps, reliability, and security best practices.
  • Collaborate cross-functionally with engineering, data, QA, security, and IT.
  • Mentor engineers and contribute to technical leadership.

Skills

SRE/DevOps experience
Cloud platforms (AWS/GCP/Azure)
Infrastructure as code (Terraform, CFN
CI/CD pipelines (GitLab, GitHub, Jira)
Kubernetes multi-cluster
Observability tooling (Prometheus, Dat
Incident response & RCA
Security & compliance
Automation scripting (Python/Go/Bash)
Networking fundamentals
Regulated environments (ISO 27001/NIST

Education

BS in Computer Science or related field
Equivalent experience

Tools

Terraform
CloudFormation
Ansible
Kubernetes
GitLab CI
GitHub Actions
Jenkins
ArgoCD
Flux

Job description

GRAIL is seeking a Staff Site Reliability / DevOps Engineer to lead the reliability, scalability, and security of our cloud-native platform. You will define infrastructure standards, drive reliability decisions, and mentor teams while building scalable systems.

Onsite at Menlo Park, CA, with planned Fall 2026 move to Sunnyvale, CA. You will own Kubernetes reliability, establish SLO/SLI, and lead incident response across AWS/GCP/Azure environments.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Staff Site Reliability Engineer (SRE) | Dev Ops Engineer #4770
Staff Site Reliability Engineer (SRE) | Dev Ops Engineer #4770

Grailbio • Edison (CA)

On-site
USD 169,000 - 224,000
Flexible time-off
401(k) with employer match
Medical, dental, vision coverage
Senior Cloud Security Engineer - AWS & DevSecOps
Senior Cloud Security Engineer - AWS & DevSecOps

Initial Therapeutics, Inc. • Menlo Park (CA)

Hybrid
USD 169,000 - 224,000
Senior Staff SRE: Global Reliability & Platform Architect
Senior Staff SRE: Global Reliability & Platform Architect

LiveRamp • San Francisco (CA)

On-site
USD 181,000 - 263,000
Flexible paid time off
Comprehensive benefits package
401K matching plan
Staff SRE: Platform Reliability Architect
Staff SRE: Platform Reliability Architect

Anduril Industries • Costa Mesa (CA)

On-site
USD 191,000 - 253,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Gruve • Redwood City (CA)

On-site
USD 120,000 - 150,000
Staff SRE: Release Engineering & Reliability
Staff SRE: Release Engineering & Reliability

Plaid • San Francisco (CA)

On-site
USD 207,000 - 274,000
Equity
Commission
Staff SRE Engineer: AI-Driven Platform Reliability
Staff SRE Engineer: AI-Driven Platform Reliability

AI Chopping Block • Costa Mesa (CA), Northern (KY)

Hybrid
USD 191,000 - 253,000
Lead Site Reliability Engineer - Multi-Cloud & Kubernetes
Lead Site Reliability Engineer - Multi-Cloud & Kubernetes

SRI Tech Solutions Inc. • Orlando (FL)

On-site
USD 140,000 - 190,000
Senior Staff SRE - Lead Resilient, Scalable Cloud Platform
Senior Staff SRE - Lead Resilient, Scalable Cloud Platform

jobr.pro • San Francisco (CA)

Hybrid
USD 220,000 - 270,000
100% health coverage for employees
Market-leading leave policies
Paid time off
+2
Staff SRE — Platform Reliability Lead
Staff SRE — Platform Reliability Lead

United States Digital Space LLC • United States

Hybrid
USD 127,000 - 161,000
Deutschlandticket (Germany-wide public
28 vacation days
Work from abroad up to 10 days/year
+9