Data Reliability Engineer

BridgeView

United States

On-site

USD 120,000 - 180,000

Full time

2 days ago
Be an early applicant
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Job summary

BridgeView seeks a Data Reliability Engineer to own the reliability and performance of our AWS-based data platform. You will manage production data pipelines, monitor SLAs, diagnose incidents, and implement durable fixes.

Collaborate with engineering to improve design and operational practices, develop monitoring and automation, and contribute to disaster recovery planning. Occasional off-hours support may be required. 5+ years in data engineering with Python/SQL and Spark/EMR a plus.

Qualifications

  • Bachelor's degree in Computer Science, Information Systems, Data Science, or related field.
  • 5+ years in data engineering or analytics platform roles, with 3+ years operating production cloud data warehouses (Redshift, Snowflake, etc.).
  • 3+ years building AWS data pipelines and managing them through production.
  • 3+ years working with production data platforms in AWS, focusing on anomaly detection, reconciliation, and end-to-end validation.
  • 3+ years experience with Python and SQL in real data systems.
  • Hands-on experience troubleshooting distributed data processing systems such as Spark/EMR, Redshift, and streaming systems.
  • Experience with AWS data services (EMR, Redshift, DynamoDB, S3, or similar).
  • Proven ability to debug and resolve production issues in data pipelines and platforms.
  • Strong problem-solving mindset and ability to work through ambiguous production issues.

Responsibilities

  • Own the reliability and stability of production data pipelines and platform services.
  • Define and enforce data SLAs/SLOs for batch and streaming products.
  • Diagnose and resolve pipeline failures, delays, and data quality issues in production.
  • Investigate issues across distributed data systems, including Spark/EMR, ingestion pipelines, and warehouse performance.
  • Lead or support incident response, including triage, mitigation, and long-term resolution.
  • Perform root cause analysis and implement durable fixes to prevent recurrence.
  • Design and enhance monitoring, alerting, and observability for data systems.
  • Develop automation and tooling to reduce operational toil and improve resilience.
  • Contribute to disaster recovery planning, including backup validation and recovery workflows.
  • Partner with engineering teams to improve pipeline design, reliability, and readiness.
  • Create and maintain runbooks, SOPs, and operational documentation.
  • Participate in occasional off-hours support for production data systems when required.

Skills

Python
SQL
Data engineering
Incident response
Problem solving

Education

Bachelor's degree in Computer Science, Information Systems, Data Science, or related field

Tools

Spark/EMR
Redshift
S3/DynamoDB
Kafka/Kinesis

Job description

Data Reliability Engineer ensures the reliability, stability, and operational excellence of an AWS-based data platform. Owns production data pipelines, monitors SLAs, diagnoses incidents, and implements durable fixes. Collaborates with engineering teams to enhance design and operational practices.

Key Responsibilities:
  • Own the reliability and stability of production data pipelines and platform services.
  • Define and enforce data SLAs/SLOs for batch and streaming products.
  • Diagnose and resolve pipeline failures, delays, and data quality issues in production.
  • Investigate issues across distributed data systems, including Spark/EMR, ingestion pipelines, and warehouse performance.
  • Lead or support incident response, including triage, mitigation, and long-term resolution.
  • Perform root cause analysis and implement durable fixes to prevent recurrence.
  • Design and enhance monitoring, alerting, and observability for data systems.
  • Develop automation and tooling to reduce operational toil and improve resilience.
  • Contribute to disaster recovery planning, including backup validation and recovery workflows.
  • Partner with engineering teams to improve pipeline design, reliability, and readiness.
  • Create and maintain runbooks, SOPs, and operational documentation.
  • Participate in occasional off-hours support for production data systems when required.
Qualifications:
  • Bachelor's degree in Computer Science, Information Systems, Data Science, or related field.
  • 5+ years in data engineering or analytics platform roles, with 3+ years operating production cloud data warehouses (Redshift, Snowflake, etc.).
  • 3+ years building AWS data pipelines and managing them through production.
  • 3+ years working with production data platforms in AWS, focusing on anomaly detection, reconciliation, and end-to-end validation.
  • 3+ years experience with Python and SQL in real data systems.
  • Hands-on experience troubleshooting distributed data processing systems such as Spark/EMR, Redshift, and streaming systems.
  • Proven ability to debug and resolve production issues in data pipelines and platforms.
  • Experience with AWS data services (EMR, Redshift, DynamoDB, S3, or similar).
  • Proven ability in handling production incidents and performing root cause analysis.
  • Strong problem-solving mindset and ability to work through ambiguous production issues.
Preferred Skills:
  • Experience handling real-world data issues such as pipeline delays or failures.
  • Experience with data backfills and reprocessing.
  • Experience influencing or guiding data pipeline reliability and operational practices.
  • Exposure to streaming/event-driven systems (Kafka, Kinesis, CDC patterns).
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Data Engineer
Data Engineer

Compunnel, Inc. • Boston (MA)

On-site
USD 110,000 - 140,000
Data Reliability Engineer
Data Reliability Engineer

Selby Jennings • Chicago (IL)

On-site
USD 110,000 - 180,000
Data Reliability Engineer: AWS Pipelines & Observability
Data Reliability Engineer: AWS Pipelines & Observability

BridgeView • United States

On-site
USD 120,000 - 180,000
Data Engineer
Data Engineer

The Value Maximizer • South Carolina

On-site
USD 90,000 - 120,000
Data Engineer
Data Engineer

Magpie Health Analytics, Inc. • Baltimore (MD)

On-site
USD 110,000 - 165,000
Data Engineer
Data Engineer

Compunnel, Inc. • Westlake (OH)

On-site
USD 90,000 - 120,000
AWS Data Engineer
AWS Data Engineer

Qode Page • United States

On-site
USD 140,000 - 180,000
Senior Data Reliability Engineer AWS
Senior Data Reliability Engineer AWS

Koitecc Solutions • Northern (KY)

Hybrid
USD 106,000 - 149,000
Medical insurance
Dental insurance
Vision insurance
+2
Senior Data Engineer: Real-Time Pipelines on AWS
Senior Data Engineer: Real-Time Pipelines on AWS

eOne Infotech • Fort Mill (SC)

On-site
USD 90,000 - 120,000
Manager – Data Engineering
Manager – Data Engineering

Simplify Recruiting • United States

On-site
USD 150,000 - 230,000