Site Reliability Engineering (SRE) Manager

McAfee, Inc.

Frisco (TX)

On-site

USD 124,000 - 230,000

Full time

6 days ago
Be an early applicant
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Benefits offered by this job

Bonus Program
401k Retirement
Medical Insurance
Parental Leave
Paid Holidays
Unlimited PTO
Sick/ Vacation accrual

Job summary

McAfee is seeking an experienced SRE Manager to lead the North American Site Reliability Engineering team and own reliability strategy across Cloud and Kubernetes platforms. This onsite role is located in Frisco, TX, with candidates expected to be within commutable distance.

The role focuses on leading SREs, incident management, automation with Python, Terraform IaC, and observability using Grafana and CloudWatch to ensure resilient services.

Qualifications

  • 9+ years in Site Reliability Engineering, DevOps, or related roles with leadership experience.
  • Strong AWS, EKS/GKE troubleshooting and architectural guidance.
  • Certifications: AWS Solutions Architect – Professional or AWS DevOps Engineer – Professional; CKA or equivalent.

Responsibilities

  • Lead and grow the SRE team, setting direction and reliability strategy.
  • Own Incident and Problem Management processes with executive communication.
  • Drive automation strategy using Python tooling to reduce toil at scale.
  • Set standards for Terraform-based infrastructure as code across teams.
  • Define observability strategy with Grafana, alerts, and CloudWatch analysis.
  • Serve as senior escalation point and incident commander for high-severity incidents.
  • Partner with senior leadership, product, and engineering on risk and remediation roadmaps.
  • Own hiring, mentoring, and performance management for the SRE team.
  • Design and build self-healing automation and runbooks for known failure patterns.
  • Implement monitoring across multiple regions for health, latency, and failover readiness.
  • Identify bottlenecks and automate recurring tasks to reduce ops workload.
  • Drive ITSM maturity and integrate Incident/Problem Management best practices.
  • Manage on-call structure, escalation paths, and team readiness.
  • Report reliability metrics and improvement initiatives to leadership.

Skills

AWS
Python
Terraform
Kubernetes
Observability
Incident Mgmt
Leadership
Communication
SRE/DevOps
EKS/GKE

Education

Bachelor's degree in CS/IT
Master's degree or MBA

Tools

Grafana
CloudWatch

Job description

Role Summary

It's an exciting time to join McAfee!

We're looking for an experienced SRE Manager to lead our growing North American Site Reliability Engineering team and own reliability strategy across our Cloud and Kubernetes platforms. You'll combine deep technical credibility with strong people leadership — setting direction, growing the team, and acting as the senior escalation point during the most critical incidents — while partnering closely with engineering and business leadership.

This is an onsite position located in our Frisco, TX office. We are only considering candidates within a commutable distance to the Frisco office.

Position Details
About the role:
  • Lead and grow a team of SREs, setting technical direction and reliability strategy across AWS, GCP infrastructure and EKS platforms, and GKE platforms.
  • Own the organization's Incident and Problem Management processes, ensuring major incidents are handled efficiently, with timely executive communication and thorough post-incident reviews.
  • Drive the team's automation strategy, championing Python-based tooling and frameworks that reduce manual toil and improve reliability at scale.
  • Set standards for Terraform-based infrastructure-as-code, ensuring secure, scalable, and consistent provisioning practices across teams.
  • Define the organization's observability strategy, ensuring Grafana dashboards, alerting, and SQL/CloudWatch-based analysis practices scale effectively.
  • Act as a senior escalation point and incident commander for the most critical, high-severity incidents, providing calm, decisive leadership under pressure.
  • Partner with senior leadership, product, and engineering stakeholders to communicate risk, reliability posture, and remediation roadmaps clearly and confidently.
  • Own hiring, mentoring, performance management, and career development for the SRE team.
  • Design and build self-healing automation and runbooks that detect known failure patterns and trigger remediation automatically, reducing manual intervention and recovery time for recurring incidents.
  • Implement and maintain monitoring across multiple regions to ensure consistent visibility into system health, latency, and failover readiness across all deployment zones.
  • Proactively identify potential failure points and performance bottlenecks before they impact production and reduce operational workload by automating recurring manual tasks.
  • Drive ITSM process maturity across the organization, partnering with other teams to embed Incident and Problem Management best practices.
  • Manage on-call structure, escalation paths, staffing, and operational readiness for the team.
  • Report on reliability metrics, incident trends, and improvement initiatives to senior leadership.
About You:
  • 9+ years of experience in Site Reliability Engineering, DevOps, Infrastructure, or related roles, including significant experience in a leadership or management capacity.
  • Demonstrated experience building, leading, and growing high-performing technical teams.
  • Strategic reliability leadership: Balances operational excellence with long-term reliability improvements, ensuring the team addresses immediate risks while building scalable, sustainable practices.
  • Deep, hands-on background with AWS infrastructure and strong technical credibility to guide architecture and operational decisions.
  • Proven track record leading teams through complex EKS/GKE troubleshooting and operational challenges.
  • Strong technical fluency in Python for automation and Terraform for infrastructure-as-code, with the ability to guide and review the team's work.
  • Extensive experience owning ITSM processes — Incident and Problem Management — at an organizational level.
  • Strong command of observability practices, including Grafana dashboards, SQL, CloudWatch Logs Insights, and alerting strategy.
  • Outstanding communication skills — able to clearly articulate technical risk, incident impact, and strategy to executive leadership and cross-functional stakeholders.
  • Strong stakeholder management skills, comfortable operating at the intersection of engineering, product, and business leadership.
  • AWS Certification — required (e.g., AWS Certified Solutions Architect – Professional, AWS Certified DevOps Engineer – Professional, or equivalent).
  • Certified Kubernetes Administrator (CKA) or equivalent EKS/Kubernetes certification — required.
  • Incident leadership: Provides calm, decisive direction during high-severity and high-visibility incidents, helping teams stay focused and coordinated under pressure.
  • Stakeholder trust: Builds strong working relationships with cross-functional teams, engineering leaders, product partners, and executive stakeholders.
  • Degrees are a plus, including Bachelor's degree in Computer Science, Information Technology, or a related field and/or Master's degree or MBA

#LI-Onsite

Company Overview

McAfee is a leader in personal security for consumers. Focused on protecting people, not just devices, McAfee consumer solutions adapt to users’ needs in an always online world, empowering them to live securely through integrated, intuitive solutions that protects their families and communities with the right security at the right moment.

Company Benefits and Perks

We work hard to embrace diversity and inclusion and encourage everyone at McAfee to bring their authentic selves to work every day. We offer a variety of social programs, flexible work hours and family-friendly benefits to all of our employees.:

  • Bonus Program
  • 401k Retirement
  • Medical, Dental, Vision, Basic Life, Short Term Disability and Long-Term Disability Coverage
  • Paid Parental Leave
  • Support and Community Involvement
  • 14 Paid Company Holidays
  • Unlimited Paid Time Off for Exempt Employees
  • 96 Hours of Sick Time and 120 Hours of Vacation for Non-Exempt Employees Accrued Each Year

We're serious about our commitment to diversity which is why McAfee prohibits discrimination based on race, color, religion, gender, national origin, age, disability, veteran status, marital status, pregnancy, gender expression or identity, sexual orientation or any other legally protected status.

Pay Range

The anticipated compensation for this position is USD $123,650.00/Yr. - USD $229,650.00/Yr. depending on experience and qualifications.

Job Applicant Privacy Notice

Please click here to view and download the Job Applicant Privacy Notice, which applies to all McAfee job applicants who are residents of the state of California.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Site Reliability Engineering (SRE) Manager
Site Reliability Engineering (SRE) Manager

McAfee GmbH • Frisco (TX)

On-site
USD 124,000 - 230,000
Bonus Program
401k Retirement
Medical, Dental, Vision
Applied AI Engineer
Applied AI Engineer

McAfee GmbH • Frisco (TX)

Hybrid
USD 109,000 - 179,000
Bonus Program
401k Retirement
Medical, Dental, Vision, Basic Life, S
+3
Senior Platform Engineer - Frisco
Senior Platform Engineer - Frisco

McAfee • Town of Texas (WI)

Hybrid
USD 107,000 - 177,000
Bonus Program
401k Retirement
Medical, Dental, Vision
+2
Applied AI Engineer
Applied AI Engineer

McAfee • Frisco (TX)

Hybrid
USD 109,000 - 179,000
Bonus Program
401k Retirement
Benefits package
+3
Applied AI Engineer
Applied AI Engineer

McAfee • San Francisco (CA)

Hybrid
USD 109,000 - 179,000
Bonus Program
401k Retirement
Medical, Dental, Vision, Basic Life, S
+4
HR Generalist - Hybrid
HR Generalist - Hybrid

McAfee • San Jose (CA)

Hybrid
USD 95,890 - 157,540
Bonus Program
401k Retirement
Medical, Dental, Vision
+4
HR Generalist - Hybrid
HR Generalist - Hybrid

McAfee • New York (NY)

Hybrid
USD 96,000 - 158,000
Bonus Program
401k Retirement
Medical, Dental, Vision coverage
+2
Lead AI Engineer – eCommerce Engineering
Lead AI Engineer – eCommerce Engineering

McAfee • Frisco (TX)

Hybrid
USD 135,000 - 224,000
Bonus Program
401(k) Retirement
Medical, Dental, Vision coverage
+3
HR Generalist - Hybrid
HR Generalist - Hybrid

McAfee • Frisco (TX)

Hybrid
USD 96,000 - 158,000
Bonus Program
401k Retirement
Medical, Dental, Vision, Basic Life, S
+4
Sr. Manager, Retention Analytics
Sr. Manager, Retention Analytics

McAfee • San Jose (CA)

Hybrid
USD 136,000 - 223,000
Bonus Program
401k Retirement
Medical, Dental, Vision, Basic Life, S
+7