SRE Manager: Cloud & Kubernetes Reliability Lead

McAfee, Inc.

Frisco (TX)

On-site

USD 124,000 - 230,000

Full time

6 days ago
Be an early applicant
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Benefits offered by this job

Bonus Program
401k Retirement
Medical Insurance
Parental Leave
Paid Holidays
Unlimited PTO
Sick/ Vacation accrual

Job summary

McAfee is seeking an experienced SRE Manager to lead the North American Site Reliability Engineering team and own reliability strategy across Cloud and Kubernetes platforms. This onsite role is located in Frisco, TX, with candidates expected to be within commutable distance.

The role focuses on leading SREs, incident management, automation with Python, Terraform IaC, and observability using Grafana and CloudWatch to ensure resilient services.

Qualifications

  • 9+ years in Site Reliability Engineering, DevOps, or related roles with leadership experience.
  • Strong AWS, EKS/GKE troubleshooting and architectural guidance.
  • Certifications: AWS Solutions Architect – Professional or AWS DevOps Engineer – Professional; CKA or equivalent.

Responsibilities

  • Lead and grow the SRE team, setting direction and reliability strategy.
  • Own Incident and Problem Management processes with executive communication.
  • Drive automation strategy using Python tooling to reduce toil at scale.
  • Set standards for Terraform-based infrastructure as code across teams.
  • Define observability strategy with Grafana, alerts, and CloudWatch analysis.
  • Serve as senior escalation point and incident commander for high-severity incidents.
  • Partner with senior leadership, product, and engineering on risk and remediation roadmaps.
  • Own hiring, mentoring, and performance management for the SRE team.
  • Design and build self-healing automation and runbooks for known failure patterns.
  • Implement monitoring across multiple regions for health, latency, and failover readiness.
  • Identify bottlenecks and automate recurring tasks to reduce ops workload.
  • Drive ITSM maturity and integrate Incident/Problem Management best practices.
  • Manage on-call structure, escalation paths, and team readiness.
  • Report reliability metrics and improvement initiatives to leadership.

Skills

AWS
Python
Terraform
Kubernetes
Observability
Incident Mgmt
Leadership
Communication
SRE/DevOps
EKS/GKE

Education

Bachelor's degree in CS/IT
Master's degree or MBA

Tools

Grafana
CloudWatch

Job description

McAfee is seeking an experienced SRE Manager to lead the North American Site Reliability Engineering team and own reliability strategy across Cloud and Kubernetes platforms. This onsite role is located in Frisco, TX, with candidates expected to be within commutable distance.

The role focuses on leading SREs, incident management, automation with Python, Terraform IaC, and observability using Grafana and CloudWatch to ensure resilient services.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

SRE Manager: Cloud & Kubernetes Reliability Leader
SRE Manager: Cloud & Kubernetes Reliability Leader

McAfee GmbH • Frisco (TX)

On-site
USD 124,000 - 230,000
Bonus Program
401k Retirement
Medical, Dental, Vision
Site Reliability Engineering (SRE) Manager
Site Reliability Engineering (SRE) Manager

McAfee GmbH • Frisco (TX)

On-site
USD 124,000 - 230,000
Bonus Program
401k Retirement
Medical, Dental, Vision
Site Reliability Engineering (SRE) Manager
Site Reliability Engineering (SRE) Manager

McAfee, Inc. • Frisco (TX)

On-site
USD 124,000 - 230,000
Bonus Program
401k Retirement
Medical Insurance
+4
SRE Manager: Reliability Leader for Scalable Cloud
SRE Manager: Reliability Leader for Scalable Cloud

Litera • Denver (CO)

Hybrid
USD 120,000 - 160,000
Remote SRE Engineering Manager, Platform Reliability
Remote SRE Engineering Manager, Platform Reliability

Sophos Group • Northern (KY)

Hybrid
USD 153,000 - 255,000
Remote-first work model
Remote SRE Manager: Cloud Reliability & DevOps Lead
Remote SRE Manager: Cloud Reliability & DevOps Lead

Deepwatch • Tampa (FL)

Hybrid
USD 178,000 - 213,000
Medical, dental, vision insurance
Flexible Time Off
Professional development benefits
+1
SRE Manager: Lead Reliability & Observability at Scale
SRE Manager: Lead Reliability & Observability at Scale

Iac/interactivecorp • Sacramento (CA)

On-site
USD 150,000 - 210,000
Collaborative work environment
Commitment to carbon‑reduction mission
Flex schedule
+2
Senior SRE Lead: Kubernetes, Go & Cloud Reliability
Senior SRE Lead: Kubernetes, Go & Cloud Reliability

Mphasis • Dallas (TX)

On-site
USD 120,000 - 170,000
Senior SRE: Cloud, Kubernetes & 24x7 Reliability Lead
Senior SRE: Cloud, Kubernetes & 24x7 Reliability Lead

NTT DATA, Inc. • Baltimore (MD)

On-site
USD 88,000 - 110,000
Senior SRE: Cloud Native Reliability & Kubernetes
Senior SRE: Cloud Native Reliability & Kubernetes

Cisco Systems, Inc. • San Francisco (CA)

Hybrid
USD 180,000 - 240,000