Lead Site Reliability Engineer

Empower

United States

Hybrid

USD 114,000 - 165,300

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Medical, dental, vision and life
401(k) with company match
Tuition reimbursement
Paid time off + holidays
Volunteer time
BRGs for inclusion

Job summary

Hispanic Alliance for Career Enhancement is seeking a Lead Site Reliability Engineer to drive reliability across our financial services platform. You will lead SREs, define standards, and guide infrastructure strategy using AWS, Kubernetes, and Terraform.

You will set SLOs/SLIs, manage incidents, and lead disaster recovery efforts while mentoring engineers and aligning reliability work with business priorities. A strong production HPM mindset and collaboration with Security are essential.

Qualifications

  • Bachelor's degree or equivalent practical experience.
  • 7–10 years of Site Reliability Engineering experience with leadership.
  • Proven ability to lead technical teams and drive projects to completion.
  • Extensive AWS knowledge for multi-region architectures.
  • Deep Kubernetes expertise and production-scale operations.
  • Terraform for shared platforms and frameworks.
  • Production experience with Python and/or Go.
  • Observability with Datadog and Splunk at scale.
  • Strong CI/CD and enterprise pipelines.

Responsibilities

  • Lead cross-functional reliability initiatives across value streams.
  • Define and evolve SRE best practices, tools, and methodologies.
  • Architect multi-region AWS infrastructure balancing reliability, cost, and security.
  • Establish SLOs/SLIs and incident postmortems with action items.
  • Lead disaster recovery planning for critical services.
  • Build Infrastructure as Code foundations with Terraform.
  • Design production-scale Kubernetes patterns and security.
  • Establish observability standards using Datadog and Splunk.
  • Set CI/CD standards and progressive delivery at scale.
  • Lead chaos engineering and game days for reliability testing.
  • Drive FinOps initiatives to optimize cloud spend while maintaining reliability.
  • Mentor SREs and partner with leadership to align with business priorities.

Skills

AWS
Kubernetes
Terraform
Python
Go
Datadog
Splunk
CI/CD
Incident management
Security basics
Communication
Mentoring

Education

Bachelor's degree in Computer Science/IT or related field

Tools

Terraform
Kubernetes
Datadog
Splunk

Job description

Our vision for the future is based on the idea that transforming financial lives starts by giving our people the freedom to transform their own. We have a flexible work environment, and fluid career paths. We not only encourage but celebrate internal mobility. We also recognize the importance of purpose, well-being, and work-life balance. Within Empower and our communities, we work hard to create a welcoming and inclusive environment, and our associates dedicate thousands of hours to volunteering for causes that matter most to them.

Chart your own path and grow your career while helping more customers achieve financial freedom. Empower Yourself.

***Applicants must be authorized to work for any employer in the U.S. We are unable to sponsor or take over sponsorship of an employment visa at this time, including CPT/OPT.***

Lead Site Reliability Engineer

Will combine deep technical expertise with team leadership to drive reliability across Empower's financial services platform. You will lead SREs in solving complex operational challenges, establish technical standards, and advise engineering leadership on infrastructure strategy and reliability initiatives.

What you will do:
  • Lead cross-functional reliability initiatives across multiple value streams and coordinate execution across teams.
  • Define and evolve SRE best practices, tools, and methodologies across the organization.
  • Architect enterprise-scale, multi-region AWS infrastructure that balances reliability, cost, performance, and security.
  • Establish and operate SLOs, SLIs, and error budgets for critical services, using them to drive prioritization decisions.
  • Serve as incident commander for major incidents and drive postmortems that produce completed action items and organizational learning.
  • Lead disaster recovery planning for critical financial services infrastructure.
  • Build shared Infrastructure as Code foundations in Terraform (reusable modules, standards, and patterns adopted across teams).
  • Design and implement production-scale Kubernetes patterns, including multi-tenancy, security policies, and advanced scheduling.
  • Establish observability standards and strategies using Datadog and Splunk (metrics, logging, tracing, dashboards, and alerting).
  • Set CI/CD standards and patterns, including pipeline-as-code and progressive delivery at scale.
  • Lead chaos engineering, game days, and systematic reliability testing initiatives.
  • Drive FinOps initiatives to optimize cloud spend while maintaining reliability targets.
  • Lead a functional team of SREs (without direct reports) on projects and operational initiatives.
  • Mentor SREs at multiple levels through coaching, design reviews, code reviews, and training sessions.
  • Partner with Engineering, Product, and Security leadership to align reliability work with business priorities, zero-trust architecture, and compliance controls.
What you will bring:
  • Bachelor's degree in Computer Science, Information Technology, or related field (or equivalent practical experience).
  • 7 to 10 years of Site Reliability Engineering experience (or equivalent), with demonstrated technical leadership.
  • Proven ability to lead technical teams and drive complex projects to completion.
  • Expert AWS knowledge, including designing large-scale, multi-region architectures.
  • Deep Kubernetes expertise, including advanced features, security, and production-scale operations.
  • Mastery of Infrastructure as Code using Terraform, including building shared platforms and frameworks.
  • Strong software engineering background with production experience in Python and/or Go.
  • Extensive experience with observability platforms (Datadog, Splunk) and implementing monitoring at scale.
  • Deep understanding of CI/CD principles and experience implementing enterprise-grade pipelines.
  • Proven track record leading major incidents and conducting effective postmortems.
  • Strong understanding of security, networking, and infrastructure design patterns.
  • Strong communication skills with ability to explain complex technical concepts to diverse audiences.
  • Experience mentoring engineers and building technical capabilities in teams.
What will set you apart:
  • Previous technical leadership roles (Lead, Staff, or similar) in SRE or Operational Excellence.
  • Financial services industry experience with understanding of regulatory requirements.
  • Expertise in compliance frameworks (SOC 2, PCI DSS, FINRA).
  • AWS certifications (Professional level).
  • Kubernetes certifications (CKA, CKAD, CKS).
  • Experience implementing SRE at organizations with 500+ engineers.
  • Background in chaos engineering, game days, and reliability testing practices.
  • Contributions to open-source projects with demonstrated community leadership.
  • Experience with service mesh implementation and management.
  • Track record of speaking at conferences or writing technical content.
What we offer you:
  • Medical, dental, vision and life insurance
  • Retirement savings - 401(k) plan with generous company matching contributions (up to 6%), financial advisory services, potential company discretionary contribution, and a broad investment lineup
  • Tuition reimbursement up to $5,250/year
  • Business-casual environment that includes the option to wear jeans
  • Generous paid time off upon hire - including a paid time off program plus ten paid company holidays and three floating holidays each calendar year
  • Paid volunteer time - 16 hours per calendar year
  • Leave of absence programs - including paid parental leave, paid short- and long-term disability, and Family and Medical Leave (FMLA)
  • Business Resource Groups (BRGs) - BRGs facilitate inclusion and collaboration across our business internally and throughout the communities where we live, work and play. BRGs are open to all.
Base Salary Range

$114,000.00 - $165,300.00

Equal opportunity employer • Drug-free workplace

We are an equal opportunity employer with a commitment to diversity. All individuals, regardless of personal characteristics, are encouraged to apply. All qualified applicants will receive consideration for employment without regard to age (40 and over), race, color, national origin, ancestry, sex, sexual orientation, gender, gender identity, gender expression, marital status, pregnancy, religion, physical or mental disability, military or veteran status, genetic information, or any other status protected by applicable state or local law.

Remote / Hybrid Requirements

For remote and hybrid positions you will be required to provide reliable high-speed internet with a wired connection as well as a place in your home to work with limited disruption. You must have reliable connectivity from an internet service provider that is fiber, cable or DSL internet. Other necessary computer equipment, will be provided. You may be required to work in the office if you do not have an adequate home work environment and the required internet connection.

Job Posting End Date at 12:01 am on: 07-24-2026

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Lead Site Reliability Engineer
Lead Site Reliability Engineer

Empower Retirement • Greenwood Village (CO)

Hybrid
USD 114,000 - 166,000
Medical insurance
401(k) with company match
Tuition reimbursement
+3
Site Reliability Engineer
Site Reliability Engineer

Empower Retirement, LLC • Overland Park (KS)

On-site
USD 87,400 - 123,400
Medical, dental, vision and life ins."
401(k) with company match
Tuition reimbursement
+1
Site Reliability Engineer
Site Reliability Engineer

Empower • United States

Hybrid
USD 87,000 - 123,000
Medical Insurance
Dental Insurance
Vision Insurance
+7
Lead, Site Reliability Engineer
Lead, Site Reliability Engineer

CardWorks • Pittsburgh

Hybrid
USD 146,000 - 163,000
Competitive Pay
Medical, Dental, and Vision Benefits
401(k) Plan with Company Match
+1
Senior Director Software Engineering
Senior Director Software Engineering

Empower Retirement • Greenwood Village (CO)

Hybrid
USD 152,000 - 220,000
401(k)
Tuition reimbursement
Paid time off
+3
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Storm2 • Scottsdale (AZ)

Hybrid
USD 140,000 - 150,000
Competitive healthcare, dental, and vision coverage
401(k) with company match
Generous PTO and paid holidays
+1
Lead Site Reliability Engineer
Lead Site Reliability Engineer

CardWorks Servicing LLC • Pittsburgh

On-site
USD 146,000 - 163,000
Competitive base pay
Medical, Dental and Vision coverage
401(k) Plan with Company Match
+1
Lead Site Reliability Engineer (SRE) / Principal Site Reliability Engineer (SRE)
Lead Site Reliability Engineer (SRE) / Principal Site Reliability Engineer (SRE)

Mindlance • Irving (TX)

Hybrid
USD 120,000 - 160,000
Senior Data Reliability Engineer AWS
Senior Data Reliability Engineer AWS

Empower • United States

On-site
USD 106,000 - 149,000
Medical insurance
401(k) plan with company match
Tuition reimbursement
+5
Senior DevOps/SRE Engineer
Senior DevOps/SRE Engineer

SEI • Oaks (PA)

Hybrid
USD 140,000 - 170,000
Comprehensive healthcare coverage
401(k) matching
Tuition reimbursement
+1