Director Site Reliability Engineering

Webster Bank

Southington (CT)

Hybrid

USD 135,000 - 155,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Webster Bank is seeking a Director of Site Reliability Engineer to lead reliability across our AWS-based platforms and MuleSoft integrations. You will transform operational practices with automation and cloud-native design, driving proactive improvements and scalable systems.

You will partner with software engineering teams, define SLOs/SLIs, and establish robust monitoring, incident response, and release procedures.

Qualifications

  • Bachelor’s degree in related field required.
  • 5–7 years of progressive experience in designing, building, and maintaining business applications in cloud environments.
  • Experience setting up SLAs/SLOs/SLIs for critical services.

Responsibilities

  • Monitor and achieve observability across platforms.
  • Automate deployments, monitoring, and system management.
  • Define and implement reliability patterns and SLOs/SLIs.
  • Collaborate with software teams to ensure reliability in production.
  • Create and maintain runbooks and diagrams for knowledge sharing.

Skills

AWS
MuleSoft
Python
Automation
DevOps
SRE practices
GitLab
Terraform
Dynatrace

Education

Bachelor’s Degree in Arts/Sciences (BA/BS)

Tools

Dynatrace
Okta
Apigee
GitLab
Terraform
Jira
Confluence

Job description

If you’re looking for a meaningful career, you’ll find it here at Webster. Founded in 1935, our focus has always been to put people first--doing whatever we can to help individuals, families, businesses and our colleagues achieve their financial goals. As a leading commercial bank, we remain passionate about serving our clients and supporting our communities. Integrity, Collaboration, Accountability, Agility, Respect, Excellence are Webster’s values, these set us apart as a bank and as an employer.

Come join our team where you can expand your career potential, benefit from our robust development opportunities, and enjoy meaningful work!

The Director of Site Reliability Engineer is a pivotal technical leader within the Software Engineering organization, tasked with transforming how reliability, performance, and availability are achieved across our platforms. This role goes beyond maintaining systems—it reimagines and modernizes operational practices through automation, cloud-native design, and API-driven integration.You will lead initiatives that elevate our AWS cloud architecture and MuleSoft integration ecosystem, ensuring they are secure, scalable, and resilient. By applying advanced software engineering principles and site reliability practices, you will drive a cultural and technical shift toward proactive reliability, continuous improvement, and innovation.This role requires visionary thinking, deep technical expertise in AWS and MuleSoft, and a passion for driving change that results in more reliable, efficient, and future-ready systems.

What you will do
  • Monitoring and Observability: Implement and maintain tools for monitoring, logging, and tracing to gain insights into system performance and health
  • Automation: Write software and scripts to automate repetitive tasks, such as deployment, monitoring, and system management. Advocate for and lead Automation wherever possible. Ensure environments are well-managed, structured appropriately, cost effective, and synchronized as much as possible.
  • Incident Management: Respond to incidents, troubleshoot system-level issues, and perform root cause analysis to prevent recurrence
  • Reliability Engineering: Design and build reliable and scalable systems, define Service Level Objectives (SLOs) and Indicators (SLIs), and implement reliability patterns
  • Collaboration: Work closely with software developers to ensure applications are reliable and to provide feedback on performance in a production environment
  • Documentation: Create and maintain documentation, including runbooks and system diagrams, to ensure knowledge sharing and team efficiency
  • Set a high bar for reliability and availability -- and meet the bar via automation relentless improvement.
  • Improve and sustain services through rigorous development, testing and release procedures.
  • Key player during deliberations on system design, platform management, and capacity planning.
  • Have a strong 'detective' mindset on why things don't work and be among the first to offer and work on solutions.
  • Be a 'link' between technologists and business stakeholders: able to have conversations with Line of Business (LoB) and technical Agile teams to work through challenges.
  • Partner with peers to advance the maturity of the DevOps practice including new/existing technologies, tools, processes, and standards. Clearly communicate expectations on technical direction and provide ongoing guidance.
  • Serve as a sounding board and technical advisor for your team in the analysis, design, and execution of solutions. Help your team anticipate unforeseen dependencies or gaps early in the SDLC.
  • From your domain’s viewpoint, provide leadership and technical expertise to your Agile team to validate story points are sized appropriately, sprint plans are achievable, and releases are well-planned.
  • Shared accountability with peers to ensure quality, performance, and security of systems are optimal and meet both customer SLA’s and internal/external audit expectations. Contribute and/or support others in the timely remediation of security remediation, audit, or production support issues escalated to the software engineering group. Occasional evening or weekend involvement may be needed for business-critical situations.
Skills and Abilities
  • Deep understanding of systems development life cycle, cloud-based systems, and application architecture.
  • Experience with multiple programming languages (Python, etc.), configuration management tools, and containers strongly preferred.
  • Moderate to advanced proficiency of cloud development, AWS services (ECS, EKS, S3, RDS, VPC), identity management (Okta), authorization frameworks (OAuth2), monitoring (Dynatrace), Agile DevOps (GitLab, Terraform), application security (Owasp, Veracode, AppScan), and API/Integrations (Apigee, MuleSoft, BizTalk).
  • Familiarity with unit testing concepts and test automation frameworks (SpecFlow, SOAPUI), RESTful APIs and micro-services, WCF services, TSQL, SQL queries and stored procedures.
  • Advanced knowledge and strong desire to work with Agile methodology required (SAFe, Scrum) and experience with Agile tools (Jira, Confluence).
Education Qualifications
  • Bachelor’s Degree in Arts/Sciences (BA/BS) in related field required
Experience Qualifications
  • 5-7 years of progressive working experience in designing, building, and maintaining business applications across systems and networks of moderate to high complexity in cloud hosted environments required
  • Experience in setting up SLAs/SLOs/SLIs for critical services and establishing monitoring required

    The estimated salary range for this position is $135,000.00 to $155,000.00. Actual salary may vary up or down depending on job-related factors which may include knowledge, skills, experience, and location. In addition, this position is eligible for incentive compensation.

    #LI-Hybrid

    #LI-FO1

    Webster Financial Corporation and its subsidiaries (“Webster”) are equal opportunity employers that are committed to sustaining an inclusive environment. All qualified applicants will receive consideration for employment without regard to race, color, religion, age, marital status, national origin, ancestry, citizenship, sex, sexual orientation, gender identity and/or expression, physical or mental disability, protected veteran status, or any other characteristic protected by law.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Director Site Reliability Engineering
Director Site Reliability Engineering

Webster Bank group • Stamford (CT)

Hybrid
USD 135,000 - 155,000
Senior Cloud Reliability Lead (AWS + MuleSoft)
Senior Cloud Reliability Lead (AWS + MuleSoft)

Webster Bank • Southington (CT)

Hybrid
USD 135,000 - 155,000
MD Product Owner Banking System (Online Commercial)
MD Product Owner Banking System (Online Commercial)

Webster Bank • Stamford (CT)

On-site
USD 170,000 - 196,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Jobgether • United States

Hybrid
USD 120,000 - 160,000
Competitive compensation package
Flexible work arrangements
Professional development opportunities
+2
Lead Site Reliability Engineer, Vice President
Lead Site Reliability Engineer, Vice President

Socket.dev • New York (NY)

On-site
USD 150,000 - 190,000
Site Reliability Engineer
Site Reliability Engineer

PowerToFly • Alpharetta (GA)

On-site
USD 100,000 - 130,000
Director Quant Engineer, Finance Analytics
Director Quant Engineer, Finance Analytics

Webster Bank • Stamford (CT)

On-site
USD 160,000 - 180,000
Lead Site Reliability Engineer
Lead Site Reliability Engineer

CardWorks Servicing LLC • Pittsburgh

On-site
USD 146,000 - 163,000
Competitive base pay
Medical, Dental and Vision coverage
401(k) Plan with Company Match
+1
Sr. SRE - Site Reliability Engineer
Sr. SRE - Site Reliability Engineer

Charles Schwab Corporation • Austin (TX)

On-site
USD 140,000 - 190,000
Sr. Site Reliability Engineer
Sr. Site Reliability Engineer

State of Wisconsin Investment Board • Madison (WI)

On-site
USD 150,000 - 190,000