Remote Principal SRE Team Lead | Global Cloud Reliability

Jobgether

United States

Hybrid

USD 170,000 - 210,000

Full time

44 hours ago
Be an early applicant
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Benefits offered by this job

Annual bonus
Medical, dental, and vision insurance
Life and disability insurance
Paid time off and holidays
Remote and/or hybrid work options
Equity awards for eligible positions

Job summary

Partner Company in the United States is seeking a Principal Site Reliability Engineer to own reliability, availability, and operational health of a global cloud-native AI platform. You will combine hands-on SRE work with technical leadership to shape strategy across production services.

Responsibilities include setting SLI/SLO/SLA governance, observability, incident response, and production readiness, while mentoring engineers and guiding architecture reviews across distributed teams.

Qualifications

  • 8+ years in Site Reliability Engineering, DevOps, or cloud-related roles with leadership exposure.
  • Ability to set technical direction and influence teams with authority.
  • Hands-on with container tech: Kubernetes, Docker, Istio.
  • Strong cloud experience: Azure preferred, AWS/GCP exposure.
  • Experience with observability: metrics, dashboards, alerting (Prometheus, Grafana, Loki/Thanos).
  • CI/CD and IaC: Terraform, Flux; scripting in Python/Go/Shell.

Responsibilities

  • Provide technical leadership across the SRE function and mentor engineers across locations.
  • Define reliability roadmap for 2-3 quarters; govern SLI/SLO/SLA frameworks.
  • Design sustainable on-call model; manage incident responses and postmortems.
  • Lead Production Readiness and non-functional reviews with development teams.
  • Drive CI/CD automation and collaborate with DevOps/platform teams to evolve infrastructure.
  • Communicate reliability posture to technical and non-technical stakeholders.

Skills

SRE leadership
Technical leadership
Kubernetes
Docker
Istio
Azure
AWS
GCP
Observability
Prometheus
Grafana
Terraform
Flux
Python/Go/Shell
UNIX/Linux
Networking fundamentals
English communication
Incident management

Tools

Kubernetes
Docker
Istio
Azure
AWS
GCP
Zabbix
Prometheus
Grafana
Terraform
Flux
Loki
Thanos
Jira
Confluence

Job description

Partner Company in the United States is seeking a Principal Site Reliability Engineer to own reliability, availability, and operational health of a global cloud-native AI platform. You will combine hands-on SRE work with technical leadership to shape strategy across production services.

Responsibilities include setting SLI/SLO/SLA governance, observability, incident response, and production readiness, while mentoring engineers and guiding architecture reviews across distributed teams.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Site Reliability Engineering Team Lead (Principal SRE)
Site Reliability Engineering Team Lead (Principal SRE)

Jobgether • United States

Hybrid
USD 170,000 - 210,000
Annual bonus
Medical, dental, and vision insurance
Life and disability insurance
+3
Remote Principal SRE - AI-Driven Reliability Leader
Remote Principal SRE - AI-Driven Reliability Leader

Tandem Diabetes Care, Inc. • Northern (KY)

Hybrid
USD 165,000 - 185,000
Remote Principal SRE: AI-Driven Reliability Leader
Remote Principal SRE: AI-Driven Reliability Leader

Optum • Eden Prairie (MN)

On-site
USD 135,000 - 231,000
Remote work within the U.S.
Office in Minneapolis or DC four days/
Senior SRE Leader: Cloud, Reliability & Scale
Senior SRE Leader: Cloud, Reliability & Scale

AVG • Tempe (AZ), Northern (KY)

Hybrid
USD 140,000 - 190,000
Principal SRE Lead — Cloud Reliability Architect
Principal SRE Lead — Cloud Reliability Architect

Cerence AI • United States

Hybrid
USD 180,000 - 260,000
Annual bonus
Insurance coverage
Paid time off
+4
Remote Principal SRE — AI-Driven Reliability Leader
Remote Principal SRE — AI-Driven Reliability Leader

Optum • Eden Prairie (MN)

Remote
USD 135,000 - 231,000
Senior SRE Leader: AI-Powered Reliability for Cloud
Senior SRE Leader: AI-Powered Reliability for Cloud

Optum • Minnetonka (MN)

On-site
USD 120,000 - 160,000
Comprehensive benefits package
Incentive and recognition programs
401k contribution
Remote SRE Manager: Lead AI-Driven Reliability & Cloud Ops
Remote SRE Manager: Lead AI-Driven Reliability & Cloud Ops

Arcoro Holdings Corp • Phoenix (AZ), Northern (KY)

Hybrid
USD 200,000 - 220,000
Remote Work
401(k) with Company match
Flexible PTO and Company-paid holidays
Senior SRE & Cloud Reliability Architect
Senior SRE & Cloud Reliability Architect

United States Digital Space LLC • United States

Remote
USD 180,000 - 250,000
Remote Principal SRE: AI-Driven Reliability Leader
Remote Principal SRE: AI-Driven Reliability Leader

Tandem Diabetes Care • United States

Remote
USD 165,000 - 185,000
Health insurance
401(k) with company match
Employee Stock Purchase plan
+1