Senior Manager - Site Reliability Engineering (SRE)

Ferguson Enterprises, Inc.

United States

Hybrid

USD 106,000 - 185,000

Full time

2 days ago
Be an early applicant
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Job summary

Ferguson Enterprises, LLC in a remote/hybrid setup is seeking a Senior Manager, Site Reliability Engineering to lead a team responsible for the stability, automation, and observability of Ferguson's digital commerce platform.

The role requires deep Linux experience, strong Kubernetes and Docker skills, and a proven track record building scalable CI/CD pipelines, IaC using Terraform, and robust incident response processes.

Qualifications

  • 8+ years administering Linux in production.
  • 3+ years leading or managing SRE/engineering teams.
  • Strong Kubernetes and Docker experience.
  • CI/CD with GitHub Actions; Terraform for IaC.
  • Puppet configuration management; Python automation.
  • Load balancing and virtualization with Nginx/Apache Tomcat.
  • Observability with Datadog and artifact management with JFrog Artifactory.
  • Incident response, on-call processes; SLIs/SLOs.
  • Experience in Agile delivery and cross-team collaboration.

Responsibilities

  • Lead and develop a high-performing group of SREs.
  • Set technical vision for SRE domain; drive Kubernetes adoption.
  • Improve CI/CD standards; pipelines with GitHub Actions; IaC with Terraform.
  • Define config management using Puppet and Python.
  • Guide load balancing and virtualization using F5/VMware; tune Nginx/Apache Tomcat.
  • Champion observability with Datadog and Artifactory.
  • Own incident management and on-call; conduct RCAs.
  • Drive SRE/DevOps principles such as SLIs/SLOs.
  • Recruit, mentor and retain SRE talent.
  • Partner with Product, Architecture, Security, Networking.

Skills

Linux administration
Team leadership
Kubernetes
Docker
GitHub Actions
Terraform
Puppet
Python
Nginx
Apache Tomcat
Datadog
JFrog Artifactory
Incident response
SRE/DevOps principles
Agile

Tools

GitHub Actions
Terraform
Puppet
Datadog
JFrog Artifactory
Nginx
Apache Tomcat

Job description

Job Posting: Since 1953, Ferguson has been a source of quality supplies for a variety of industries. Together We Build Better infrastructure, better homes and better businesses. We exist to make our customers’ complex projects simple, successful, and sustainable. We proactively solve problems, adapt and grow to continuously serve our customers, communities and each other. Ferguson, a Fortune 500 company, is proud to provide best-in-class products, service and capabilities across the following industries: Commercial/Mechanical, Facilities Supply, Fire and Fabrication, HVAC, Industrial, Residential Trade, Residential Building and Remodel, Waterworks and Residential Digital Commerce. Ferguson has approximately 36,000 associates across 1,700 locations. Ferguson is a community of proud associates who operate with the shared purpose of building something meaningful. You will build a career that you are proud of, at a company you can believe in.

Senior Manager - Site Reliability Engineering

Join Ferguson as a Senior Manager, Site Reliability Engineering, and help shape the reliability, performance, and availability of our digital commerce platform. In this leadership role, you’ll guide a team of Site Reliability Engineers responsible for the stability, automation, and observability of the Linux-based systems and infrastructure that power Ferguson's digital commerce, while driving the modernization of our SRE and DevOps practices. You’ll partner closely with Backend Engineers, Frontend Engineers, Principal Solution Architects, Infrastructure Engineers, and Networking teams to deliver resilient platform capabilities, improve engineering standards, and pull digital commerce forward into more modern standard processes. This is an opportunity for a proven, systems-thinking SRE leader who is passionate about developing high-performing teams, driving engineering excellence, and delivering meaningful business outcomes through reliability engineering and automation. The ideal candidate is technically strong enough to set strategic direction for the SRE domain — at least at the level of a Lead Site Reliability Engineer — while coaching senior and lead engineers to raise the reliability bar across the platform. Location: Remote (Eastern or Central Time Zone) or Hybrid (Newport News, VA), in accordance with company policy. Applicants must be based within the Eastern or Central time zones.

Duties and Responsibilities

Lead and develop a high-performing group of Site Reliability Engineers who coordinate the reliability, availability, and performance of Linux-based systems and infrastructure supporting digital commerce applications.

Set technical and strategic vision for the SRE domain, driving adoption of Kubernetes-based container platforms, Docker workloads, and modern SRE/DevOps practices across teams.

Establish and continuously improve CI/CD standards, driving improvements to pipelines built with GitHub Action Workflows and other modern tooling, including infrastructure-as-code practices using Terraform.

Define configuration management and automation standards using Puppet and Python to reduce toil and eliminate repetitive manual work at scale.

Guide strategy for load balancing, application delivery, and virtualized infrastructure using F5 and VMware, and for tuning web and application servers such as Nginx and Apache Tomcat.

Champion observability using Datadog (logging, metrics, distributed traces, alerting) and sound artifact management practices using JFrog Artifactory so the team can detect and resolve issues early.

Own incident management for the domain, defining incident response and on-call practices, leading root cause analysis for significant incidents, and converting findings into durable engineering improvements.

Drive adoption of SRE/DevOps principles such as SLIs/SLOs and error budgets, balancing reliability, velocity, and technical debt across the platform.

Champion a culture of accountability, collaboration, innovation, and continuous learning, and act as a technical escalation and decision point for sophisticated reliability and infrastructure challenges.

Recruit, mentor, and retain top SRE talent while supporting career development, performance management, and succession planning, and coach lead and senior engineers to strengthen technical capability and leadership.

Partner with Product Management, Architecture, Infrastructure, Security, and Networking team members to align reliability strategy with platform and business priorities.

Qualifications and Requirements

8+ years of professional experience administering Linux systems in a production environment, including 3+ years leading or managing engineering or SRE teams. Proven success setting technical direction for reliability, automation, or infrastructure initiatives at a scope comparable to or greater than a Lead Site Reliability Engineer.

Strong hands‑on experience with Kubernetes and Docker or similar container orchestration/containerization platforms. Good experience building, maintaining, and defining standards for CI/CD pipelines using GitHub Actions or similar tooling, and solid experience with Terraform or other infrastructure‑as‑code tools.

Extensive experience with configuration management tools such as Puppet, along with extensive experience standardizing automation using Python or a similar language.

Proven expertise in load balancing/application delivery and virtualization technologies, coupled with tuning web and application servers such as Nginx and Apache Tomcat.

Good experience with observability tooling such as Datadog, including the ability to define observability standards, and with artifact repository tools such as JFrog Artifactory.

Good understanding of networking, storage, and IT infrastructure fundamentals, including security concepts relevant to enterprise infrastructure.

Experience defining incident response processes and on‑call rotation structures, and championing SRE/DevOps principles such as SLIs/SLOs and error budgets across a team.

Strong troubleshooting and systems‑thinking skills with the ability to convert incident findings into durable engineering improvements.

Experience working within Agile product delivery organizations and collaborating with Product Management, Architecture, Infrastructure, Security, and Networking team members.

Good communication, leadership, and team development skills, with a passion for mentoring engineers and building high‑performing teams.

Leadership Skills

Strong leader who builds trust, accountability, and ownership across an SRE team. Demonstrable ability to lead engineering teams through modernization of SRE/DevOps practices and organizational change.

Good communication skills coupled with the capacity to convey reliability and infrastructure concepts in terms of business value. Passion for mentoring engineers and developing future technical and reliability leaders.

Ability to establish a culture focused on engineering excellence, operational reliability, continuous improvement, and customer outcomes.

Demonstrated ability to influence without authority and collaborate across engineering, infrastructure, and networking organizations.

Strong execution mentality with a focus on delivering measurable reliability and business results while continuously improving engineering quality.

At Ferguson, we care for each other. We value our well‑being just as much as our hard work. We are committed to a holistic approach towards benefits plans and programs that support the mental, physical and financial well‑being of our associates. Our competitive offering not only includes benefits like health, dental, vision, paid time off, life insurance and a 401(k) with a company match, but our associates also enjoy additional meaningful and inclusive enhancements that are adaptable to their diverse situations and needs, including mental health coverage, gender affirming and family building benefits, paid parental leave, associate discounts, community involvement opportunities and more!

#LI-REMOTE - Pay Range: - Actual pay rate may vary depending upon location. The estimated pay range for this position is below. The specific rate will depend on a candidate’s qualifications and prior experience.

  • $9,458.97
  • $16,551.03
  • Estimated Ranges displayed are Monthly for Salaried roles OR Hourly for all other roles.
  • This role is Bonus or Incentive Plan eligible.
  • Ferguson complies with all wage regulations. The starting wage may be higher in certain locations based on local or state wage requirements.
  • The Company is an equal opportunity employer as well as a government contractor that shall abide by the requirements of 41 CFR 60-300.5(a), which prohibits discrimination against qualified protected Veterans and the requirements of 41 CFR 60-741.5(A), which prohibits discrimination against qualified individuals on the basis of disability.
  • Ferguson Enterprises, LLC. is an equal employment employer F/M/Disability/Vet/Sexual Orientation/Gender Identity.
  • Equal Employment Opportunity and Reasonable Accommodation Information Ferguson is a project success company providing expertise, solutions and products from infrastructure, plumbing and appliances to HVAC, fire, fabrication and more. As a leading value-added distributor of residential and commercial plumbing supplies and pipe, valves and fittings in the U.S., we exist to make our customers’ complex projects simple, successful and sustainable. The professionals we serve help transform the world we live in, and we are their trusted partners with the scale to provide peace of mind. Founded in 1953, Ferguson is part of Ferguson plc, which is listed on the New York Stock Exchange (NYSE: FERG) and London Stock Exchange (LSE: FERG). With approximately 36,000 associates across 1,700 locations, Ferguson plc serves customers in all 50 states, Canada, Puerto Rico, Mexico and the Caribbean.
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior Manager - Site Reliability Engineering (SRE)
Senior Manager - Site Reliability Engineering (SRE)

Ferguson Enterprises, Inc. • Newport News (VA), Northern (KY)

Hybrid
USD 113,000 - 198,000
Health insurance
Dental insurance
Vision insurance
+4
Senior Manager - Digital Engineering
Senior Manager - Digital Engineering

Ferguson Enterprises, Inc. • Newport News (VA)

Hybrid
USD 114,000 - 199,000
Health insurance
Dental insurance
Vision insurance
+3
Lead Product Owner - Data Analytics
Lead Product Owner - Data Analytics

Ferguson • United States

Hybrid
USD 95,000 - 166,000
Senior Manager - Security Engineering
Senior Manager - Security Engineering

Ferguson • United States

Remote
USD 106,000 - 185,000
Health, dental, vision
401(k) with company match
Mental health coverage
Lead Product Owner - Data Analytics
Lead Product Owner - Data Analytics

Ferguson Enterprises, LLC • United States

Hybrid
USD 95,000 - 166,000
Health benefits
401(k) with company match
Parental leave
Information Security Manager - Network Security
Information Security Manager - Network Security

Ferguson Enterprises, Inc. • Newport News (VA)

On-site
USD 108,000 - 173,000
Health insurance
Dental insurance
Vision insurance
+5
Digital Success Manager
Digital Success Manager

Ferguson Enterprises, Inc. • Northern (KY)

On-site
USD 99,000 - 158,000
Health benefits
401(k) with company match
Paid time off
Senior Internal Auditor
Senior Internal Auditor

Ferguson Enterprises, Inc. • Northern (KY)

Hybrid
USD 77,000 - 122,000
Health benefits
401(k) with company match
Paid time off
+1
Integration Specialist - Municipal Metering (Remote California)
Integration Specialist - Municipal Metering (Remote California)

Ferguson Enterprises, Inc. • Northern (KY)

Hybrid
USD 38,000 - 94,000
Health insurance
Dental insurance
Vision coverage
+7
Inside Sales Representative
Inside Sales Representative

Ferguson Enterprises, Inc. • Burlington (MA)

On-site
USD 29,000 - 47,000
Health insurance
401(k) with company match
Paid time off
+1