Data Platform SRE Engineer - Observability & Automation

Empower LLC

Austin, Northern (TX, KY)

Hybrid

USD 106,000 - 149,000

Full time

43 hours ago
Be an early applicant
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Benefits offered by this job

Medical insurance
Dental
Vision
401(k) plan
Tuition reimbursement
Paid time off
Volunteer time
BRGs

Job summary

Empower LLC is seeking a Senior Site Reliability Engineer - Data Platforms to strengthen reliability of cloud-based data pipelines. You will collaborate with Data Engineering, Cloud, Platform, Security, and FinOps teams to improve health, observability, and efficiency of production systems.

You will design Datadog/Splunk dashboards, automate tasks with Python, and apply SRE practices including SLOs, incident management, and IaC using Terraform.

Qualifications

  • 5+ years of hands-on AWS experience in production.
  • Experience in Site Reliability Engineering, Production Engineering, Platform Engineering, or Cloud Reliability Engineering.
  • Strong understanding of SRE principles and production operations.
  • Experience with Amazon Redshift or similar data platforms in production.
  • Hands-on with Datadog and Splunk for observability.
  • Strong Python programming for automation and tooling.
  • Proficient in SQL for production troubleshooting.
  • Good knowledge of Terraform and IaC, and collaboration with Cloud/Platform teams.
  • Experience with incident management, RCA, and preventive actions.

Responsibilities

  • Own and improve reliability, availability, performance, and health of production data platforms and pipelines.
  • Monitor pipeline execution, dependencies, failures, and recovery with downstream impact analysis.
  • Troubleshoot complex production issues across AWS, data platforms, and services.
  • Design and improve observability using Datadog and Splunk, including dashboards and alerts.
  • Reduce alert noise, strengthen incident detection, and contribute to post-incident reviews.
  • Participate in on-call rotation and incident response, with root-cause analysis and corrective actions.
  • Automate operational tasks and health checks using Python.
  • Apply SRE practices: SLOs, incident management, capacity planning, resilience, and DR.
  • Use Terraform/IaC for infrastructure changes and collaboration with Cloud/Platform teams.
  • Identify recurring issues and drive sustainable engineering improvements.
  • Contribute to cost-efficiency improvements across AWS and data-platform services.
  • Develop runbooks and recovery procedures.

Skills

AWS
SRE principles
Production operations
Incident management
Python programming
SQL for production
Automation

Tools

Datadog
Splunk
Terraform
Snowflake
Redshift
Python
SQL

Job description

Empower LLC is seeking a Senior Site Reliability Engineer - Data Platforms to strengthen reliability of cloud-based data pipelines. You will collaborate with Data Engineering, Cloud, Platform, Security, and FinOps teams to improve health, observability, and efficiency of production systems.

You will design Datadog/Splunk dashboards, automate tasks with Python, and apply SRE practices including SLOs, incident management, and IaC using Terraform.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

SRE Platform Engineer II: Scale, Automate Cloud Systems
SRE Platform Engineer II: Scale, Automate Cloud Systems

DAT Freight & Analytics • Portland (OR)

Hybrid
USD 95,000 - 134,000
Medical, Dental, Vision
Parental Leave
Flexible Vacation Time
+2
Site Reliability Engineer II — Scale, Automate & Observe
Site Reliability Engineer II — Scale, Automate & Observe

Worky • Denver (CO)

Hybrid
USD 95,000 - 134,000
Medical, Dental, Vision
401k matching
Employee Stock Purchase Plan
+3
Observability & DevOps Advocate
Observability & DevOps Advocate

Datadog • California (MO)

On-site
USD 130,000 - 180,000
Stock equity (RSUs)
Employee stock purchase plan (ESPP)
Career development opportunities
+3
Senior SRE: Automate Reliability & Observability
Senior SRE: Automate Reliability & Observability

United States Digital Space LLC • Charlotte (TX)

On-site
USD 153,000 - 192,000
Discretionary incentive eligible
Benefits package
Senior SRE - Cloud Reliability & Observability (Remote)
Senior SRE - Cloud Reliability & Observability (Remote)

Bazeta • Northern (KY)

Hybrid
USD 160,000 - 180,000
Hybrid work flexibility
Competitive compensation
Health benefits and PTO
+1
DataOps & Reliability Engineer
DataOps & Reliability Engineer

Compunnel, Inc. • Erie

On-site
USD 110,000 - 160,000
Senior Observability Platform Engineer - Splunk & ELK
Senior Observability Platform Engineer - Splunk & ELK

Tata Consultancy Services • San Jose (CA)

On-site
USD 94,000 - 130,000
Senior Software Engineer – Observability Platform
Senior Software Engineer – Observability Platform

Feedinkoo • United States

Remote
USD 140,000 - 220,000
401(k)
Vision insurance
Disability insurance
+2
Datadog Observability Platform Architect
Datadog Observability Platform Architect

LGBT Great Careers • New York (NY)

Hybrid
USD 110,000 - 150,000
Senior SRE: Observability, Automation & Scalable Systems
Senior SRE: Observability, Automation & Scalable Systems

Replit • Northern (KY)

Hybrid
USD 140,000 - 190,000
Competitive Salary & Equity
401(k) 4% match (US)
Health, Dental, Vision & Life
+7