SRE Manager: Cloud & Kubernetes Reliability Leader

McAfee GmbH

Frisco (TX)

On-site

USD 124,000 - 230,000

Full time

7 days ago
Be an early applicant
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Benefits offered by this job

Bonus Program
401k Retirement
Medical, Dental, Vision

Job summary

McAfee is seeking an experienced SRE Manager to lead the North American Site Reliability Engineering team from our Frisco, TX office. You will drive reliability strategy across AWS, GCP, EKS/GKE platforms and own incident management with executive communication.

You will build automation with Python, enforce IaC via Terraform, and shape observability with Grafana and CloudWatch. This role requires leadership, broad technical credibility, and a track record of delivery in high-visibility

Qualifications

  • 9+ years in Site Reliability Engineering, DevOps, Infrastructure, or related roles with leadership experience.
  • Proven track record building, leading, and growing high-performing technical teams.
  • Strong hands-on AWS infra background guiding architecture and ops decisions.

Responsibilities

  • Lead and grow a team of SREs, setting direction and reliability strategy across cloud platforms.
  • Own Incident and Problem Management processes with executive communication and post-incident reviews.
  • Drive automation using Python-based tooling to reduce toil and improve reliability.
  • Set standards for infrastructure-as-code with Terraform across teams.
  • Define observability strategy with Grafana dashboards and CloudWatch/SQL analytics.
  • Act as senior escalation point and incident commander for high-severity incidents.
  • Partner with leadership, product, and engineering to communicate risk and remediation roadmaps.
  • Own hiring, mentoring, and career development for the SRE team.
  • Design self-healing automation and runbooks to reduce manual intervention.
  • Monitor multi-region health and improve recovery time, automating recurring tasks.

Skills

Leadership
Python
AWS
Terraform
EKS/GKE
Incident leadership
Observability
Communication
Stakeholder management
On-call management

Education

Bachelor's degree in Computer Science or related field
Master's degree or MBA (plus)

Tools

Grafana
CloudWatch
SQL

Job description

McAfee is seeking an experienced SRE Manager to lead the North American Site Reliability Engineering team from our Frisco, TX office. You will drive reliability strategy across AWS, GCP, EKS/GKE platforms and own incident management with executive communication.

You will build automation with Python, enforce IaC via Terraform, and shape observability with Grafana and CloudWatch. This role requires leadership, broad technical credibility, and a track record of delivery in high-visibility

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

SRE Manager: Cloud & Kubernetes Reliability Lead
SRE Manager: Cloud & Kubernetes Reliability Lead

McAfee, Inc. • Frisco (TX)

On-site
USD 124,000 - 230,000
Bonus Program
401k Retirement
Medical Insurance
+4
Site Reliability Engineering (SRE) Manager
Site Reliability Engineering (SRE) Manager

McAfee GmbH • Frisco (TX)

On-site
USD 124,000 - 230,000
Bonus Program
401k Retirement
Medical, Dental, Vision
Site Reliability Engineering (SRE) Manager
Site Reliability Engineering (SRE) Manager

McAfee, Inc. • Frisco (TX)

On-site
USD 124,000 - 230,000
Bonus Program
401k Retirement
Medical Insurance
+4
Remote SRE Engineering Manager, Platform Reliability
Remote SRE Engineering Manager, Platform Reliability

Sophos Group • Northern (KY)

Hybrid
USD 153,000 - 255,000
Remote-first work model
SRE Manager: Reliability Leader for Scalable Cloud
SRE Manager: Reliability Leader for Scalable Cloud

Litera • Denver (CO)

Hybrid
USD 120,000 - 160,000
Remote SRE Manager: Cloud Reliability & DevOps Lead
Remote SRE Manager: Cloud Reliability & DevOps Lead

Deepwatch • Tampa (FL)

Hybrid
USD 178,000 - 213,000
Medical, dental, vision insurance
Flexible Time Off
Professional development benefits
+1
SRE Lead: Incident Commander for Cloud Reliability
SRE Lead: Incident Commander for Cloud Reliability

Us Bank • Atlanta (GA)

On-site
USD 112,000 - 131,000
Healthcare
Life Insurance
Disability
+6
Senior SRE Lead: Kubernetes, Go & Cloud Reliability
Senior SRE Lead: Kubernetes, Go & Cloud Reliability

Mphasis • Dallas (TX)

On-site
USD 120,000 - 170,000
Senior SRE: Cloud, Kubernetes & Automation
Senior SRE: Cloud, Kubernetes & Automation

Socure • Carson City (NV)

On-site
USD 150,000 - 190,000
SRE Architecture Lead: Reliability & Cloud Platform
SRE Architecture Lead: Reliability & Cloud Platform

MACHINE LEARNING TECHNOLOGIES LLC • Atlanta (GA)

On-site
USD 140,000 - 190,000