Senior Data Reliability Engineer AWS

Koitecc Solutions

Northern (KY)

Hybrid

USD 106,000 - 149,000

Full time

38 hours ago
Be an early applicant
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Benefits offered by this job

Medical insurance
Dental insurance
Vision insurance
Life insurance
401(k) with matching

Job summary

Empower is seeking a hands-on Senior Data Reliability Engineer to own the reliability, stability, and operational excellence of our AWS-based data platform. You will diagnose issues, resolve production incidents, and influence better design and practices across the data ecosystem.

You will work closely with data and platform engineering teams to ensure data pipelines meet SLAs, with strong emphasis on monitoring, automation, and disaster recovery.

Qualifications

  • Minimum 5 years of experience working with production data platforms in AWS environments
  • Experience building data pipelines through production including real-world failures and operational challenges
  • Strong experience with Python and SQL in real data systems
  • Hands-on experience troubleshooting distributed data processing systems (e.g., Spark/EMR, Redshift, streaming systems)
  • Proven ability to debug and resolve production issues in data pipelines and data platforms
  • Experience with AWS data services (EMR, Redshift, DynamoDB, S3 or similar)
  • Experience handling production incidents and performing root cause analysis
  • Strong problem-solving mindset and ability to work through ambiguous production issues

Responsibilities

  • Own the reliability and stability of production data pipelines and data platform services
  • Diagnose and resolve data pipeline failures, delays, and data quality issues in production environments
  • Investigate issues across distributed data systems (e.g., Spark/EMR workloads, ingestion pipelines, warehouse performance)
  • Lead or support incident response, including triage, mitigation, and long-term resolution
  • Perform root cause analysis (RCA) and implement durable fixes to prevent recurrence
  • Define and improve data SLAs (freshness, latency, completeness) and ensure adherence
  • Design and enhance monitoring, alerting, and observability for data systems
  • Develop automation and tooling to reduce operational toil and improve system resilience
  • Contribute to disaster recovery (DR) and resiliency planning, including backup validation and recovery workflows
  • Partner with engineering teams to improve pipeline design, reliability, and operational readiness
  • Create and maintain runbooks, SOPs, and operational documentation
  • Participate in occasional off-hours support for production data systems when required

Skills

Python
SQL
AWS
Spark/EMR
Redshift
Data pipelines
Incident management
Observability

Tools

EMR
Redshift
DynamoDB
S3
Kafka/Kinesis

Job description

Our vision for the future is based on the idea that transforming financial lives starts by giving our people the freedom to transform their own. We have a flexible work environment, and fluid career paths. We not only encourage but celebrate internal mobility. We also recognize the importance of purpose, well-being, and work-life balance. Within Empower and our communities, we work hard to create a welcoming and inclusive environment, and our associates dedicate thousands of hours to volunteering for causes that matter most to them.

Chart your own path and grow your career while helping more customers achieve financial freedom. Empower Yourself.

Applicants must be authorized to work for any employer in the U.S. We are unable to sponsor or take over sponsorship of an employment visa at this time, including CPT/OPT.

We are looking for a hands-on Senior Data Reliability Engineer to own the reliability, stability, and operational excellence of our AWS-based data platform.

This role is focused on operating, troubleshooting, and improving production data systems, ensuring that data pipelines and analytics platforms are resilient, performant, and meet business-critical SLAs.

You will work closely with data and platform engineering teams to diagnose issues, resolve production incidents, and influence better design and operational practices across the data ecosystem.

What You Will Do
  • Own the reliability and stability of production data pipelines and data platform services
  • Diagnose and resolve data pipeline failures, delays, and data quality issues in production environments
  • Investigate issues across distributed data systems (e.g., Spark/EMR workloads, ingestion pipelines, warehouse performance)
  • Lead or support incident response, including triage, mitigation, and long-term resolution
  • Perform root cause analysis (RCA) and implement durable fixes to prevent recurrence
  • Define and improve data SLAs (freshness, latency, completeness) and ensure adherence
  • Design and enhance monitoring, alerting, and observability for data systems
  • Develop automation and tooling to reduce operational toil and improve system resilience
  • Contribute to disaster recovery (DR) and resiliency planning, including backup validation and recovery workflows
  • Partner with engineering teams to improve pipeline design, reliability, and operational readiness
  • Create and maintain runbooks, SOPs, and operational documentation
  • Participate in occasional off-hours support for production data systems when required
What You Will Bring
  • Minimum 5 years of experience working with production data platforms in AWS environments
  • Prior experience building data pipelines and seeing them through production, including exposure to real-world failures and operational challenges
  • Strong experience with Python and SQL in real data systems
  • Hands-on experience troubleshooting distributed data processing systems (e.g., Spark/EMR, Redshift, streaming systems)
  • Proven ability to debug and resolve production issues in data pipelines and data platforms
  • Experience with AWS data services (such as EMR, Redshift, DynamoDB, S3, or similar)
  • Experience handling production incidents and performing root cause analysis
  • Strong problem-solving mindset and ability to work through ambiguous production issues
What Will Set You Apart
  • Experience handling real-world data issues such as pipeline delays or failures
  • Experience with backfills and reprocessing
  • Experience with late-arriving or incomplete data
  • Experience improving observability and alerting specifically for data systems
  • Experience influencing or guiding data pipeline reliability and operational practices
  • Exposure to streaming/event-driven systems (Kafka, Kinesis, CDC patterns)
  • Experience with disaster recovery, backup validation, and resiliency testing
  • Strong communication during incidents with both technical and non-technical stakeholders
Work conditions
  • Participate in an on-call rotation; occasional change windows outside business hours to support safe releases and resiliency drills.
  • This job operates in a professional office environment.

This job description is not intended to be an exhaustive list of all duties, responsibilities and qualifications of the job. The employer has the right to revise this job description at any time. You will be evaluated in part based on your performance of the responsibilities and/or tasks listed in this job description. You may be required perform other duties that are not included on this job description. The job description is not a contract for employment, and either you or the employer may terminate employment at any time, for any reason.

What we offer you

We offer an array of diverse and inclusive benefits regardless of where you are in your career. We believe that providing our employees with the means to lead healthy balanced lives results in the best possible work performance.

  • Medical, dental, vision and life insurance
  • Retirement savings - 401(k) plan with generous company matching contributions (up to 6%), financial advisory services, potential company discretionary contribution, and a broad investment lineup
  • Tuition reimbursement up to $5,250/year
  • Business-casual environment that includes the option to wear jeans
  • Generous paid time off upon hire - including a paid time off program plus ten paid company holidays and three floating holidays each calendar year
  • Paid volunteer time - 16 hours per calendar year
  • Leave of absence programs - including paid parental leave, paid short- and long-term disability, and Family and Medical Leave (FMLA)
  • Business Resource Groups (BRGs) - BRGs facilitate inclusion and collaboration across our business internally and throughout the communities where we live, work and play. BRGs are open to all.
Base Salary Range

$105,700.00 - $149,275.00

The salary range above shows the typical minimum to maximum base salary range for this position in the location listed. Non-sales positions have the opportunity to participate in a bonus program. Sales positions are eligible for sales incentives, and in some instances a bonus plan, whereby total compensation may far exceed base salary depending on individual performance. Actual compensation offered may vary from posted hiring range based upon geographic location, work experience, education, licensure requirements and/or skill level and will be finalized at the time of offer.

Equal opportunity employer Drug-free workplace

We are an equal opportunity employer with a commitment to diversity. All individuals, regardless of personal characteristics, are encouraged to apply. All qualified applicants will receive consideration for employment without regard to age (40 and over), race, color, national origin, ancestry, sex, sexual orientation, gender, gender identity, gender expression, marital status, pregnancy, religion, physical or mental disability, military or veteran status, genetic information, or any other status protected by applicable state or local law.

For remote and hybrid positions you will be required to provide reliable high-speed internet with a wired connection as well as a place in your home to work with limited disruption. You must have reliable connectivity from an internet service provider that is fiber, cable or DSL internet. Other necessary computer equipment, will be provided. You may be required to work in the office if you do not have an adequate home work environment and the required internet connection.

Job Posting End Date at 12:01 am on: 09-28-2026

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior Data Reliability Engineer AWS
Senior Data Reliability Engineer AWS

Empower • United States

Hybrid
USD 106,000 - 149,000
Medical insurance
Dental insurance
Vision insurance
+8
Data Architect
Data Architect

Empower LLC • Northern (KY)

Hybrid
USD 138,000 - 200,000
Health insurance
401(k) plan with match
Tuition reimbursement
+2
Automation Quality Engineer
Automation Quality Engineer

Empower Retirement, LLC • Kansas

Hybrid
USD 72,000 - 102,000
Medical insurance
Dental insurance
Vision insurance
+5
Architect of Data Architecture
Architect of Data Architecture

Socket.dev • Overland Park (KS)

On-site
USD 114,000 - 165,000
Medical insurance
401(k) plan
Tuition reimbursement
+1
Data engineer, Decision Intelligence Technology (AWS)
Data engineer, Decision Intelligence Technology (AWS)

Amazon • Bellevue (WA)

On-site
USD 132,000 - 179,000
Architect of Data Architecture
Architect of Data Architecture

Empower • Greenwood Village (CO)

On-site
USD 114,000 - 165,000
Medical benefits
401(k) with match
Tuition reimbursement
+2
Senior Delivery Consultant - Data , Professional Services, AWSI HCLS
Senior Delivery Consultant - Data , Professional Services, AWSI HCLS

Amazon Web Services (AWS) • Denver (CO)

On-site
USD 154,000 - 208,000
Sr. Delivery Consultant - Data , AWS Professional Services
Sr. Delivery Consultant - Data , AWS Professional Services

Amazon Web Services (AWS) • Houston (TX)

On-site
USD 154,000 - 208,000
Health insurance
401(k) matching
Paid time off
+1
Data Architect
Data Architect

Careerwebsite • Northern (KY)

Hybrid
USD 138,000 - 200,000
Medical, dental, vision and life ins.
401(k) with company match
Tuition reimbursement
+2
Senior Delivery Consultant - Data , Professional Services, AWSI HCLS
Senior Delivery Consultant - Data , Professional Services, AWSI HCLS

Amazon Web Services (AWS) • Boston (MA)

On-site
USD 154,000 - 208,000