Sr. Data Engineer

ARC-One Solutions

United States

On-site

USD 100,000 - 157,000

Full time

11 days ago
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Benefits offered by this job

Flexible work hours
Remote work with internet access
Travel as needed

Job summary

ARC-One Solutions is seeking a data engineering lead to manage and evolve an enterprise data lake and warehouse, ensuring secure and efficient data flow. The base salary range is provided, with pay determined by knowledge, skills, education, and location.

Responsibilities include designing scalable pipelines on AWS (S3, Glue, Redshift, Lambda, MWAA), developing batch and near real-time ETL/ELT workflows in Python/PySpark, and implementing governance and security controls.

Qualifications

  • Bachelor's degree in computer-related field with 5+ years in data engineering.
  • Experience building scalable AWS-based data platforms and pipelines using Lambda, Glue, Athena, S3, Redshift, DMS, MWAA, and Step Functions.
  • Advanced proficiency in Python, SQL, and PySpark with hands-on ETL/ELT framework development.
  • Experience implementing data quality, governance, and optimization practices in AWS.
  • Strong communication to translate data concepts for stakeholders; healthcare/regulatory experience preferred.
  • Experience with metadata, data lineage, observability, and data catalogs.

Responsibilities

  • Design, implement, and maintain scalable data pipelines on AWS using S3, DMS, Glue, Lambda, MWAA, and Redshift.
  • Develop robust batch and near real-time ETL/ELT workflows using Python and PySpark.
  • Implement incremental/CDC mechanisms with restart, idempotency, and recovery capabilities.
  • Introduce automated controls for completeness, accuracy, and lineage.
  • Optimize Glue/Spark, Athena, Redshift, and S3 workloads via partitioning and cost-aware design.
  • Build near real-time/event-driven pipelines with Kinesis/Kafka when needed.
  • Enforce data security, least privilege, governance, and regulatory compliance.
  • Monitor pipelines with CloudWatch; troubleshoot and resolve production incidents.
  • Promote CI/CD, Git version control, testing, and cross-team collaboration.
  • Document pipelines and operational procedures.

Skills

Python
SQL
PySpark
Data modeling
ETL/ELT design
Communication
Regulatory compliance

Education

Bachelor's degree in computer-related field

Tools

AWS Lambda
AWS Glue
Amazon Redshift
AWS S3
DMS
MWAA (Airflow)
Step Functions
Kinesis
Kafka
Git
CloudWatch

Job description

Overview

Manages and evolves the enterprise data lake and data warehouse while ensuring the reliable, secure, and efficient flow of high-quality data. Implements data processes, managing data architecture, designing ETL processes, and analyzing data for business insights.

The base salary range for this position is $99,937-$157,044.

Actual pay will be determined based upon a candidate’s job-related knowledge, skills, education, experience, geographic location, and may include other job-related factors such as certification(s), professional licensure, or internal equity considerations.

Responsibilities
  • Design, implement and maintain scalable data pipnes on WS using S3, DMS, Glue, lambda, step function/MWAA & Redshift.
  • Develop robust batch and near-real-time ETL/ELT workflow to ingest, cleanse, transform and load data from databases, legacy applications and event streams using Python & Pyspark.
  • Design incremental/CDC mechanism, including restart ability, idempotency, duplicate handling and recovery.
  • Implement automated controls for completeness, accuracy, reconciliation, schema changes & lineage.
  • Optimize Glue/Spark, Athena, Redshift & S3 workload through partitioning, columnar formats, query tuning and appropriate storage/compute design.
  • Design near real time/event-driven pipelines using Kinesis/Kafka where required, covering ordering, retry, idempotency and failure recovery.
  • Implement AWS data security, least privilege access, data classification and governance controls.
  • Monitor pipelines such as CloudWatch, troubleshoot failure and resolving production data incidents.
  • Enforce Git/version control, code review, automated testing and CI/CD practices.
  • Work with product owners, architect, reporting and business stakeholders to translate requirements into scalable data solutions.
  • Document pipelines and operational procedures.
Qualifications

Qualifications Required

  • Bachelor's degree in a computer-related field from an accredited college or university and five (5) or more years of experience in data engineering, building scalable and distributed ETL data pipelines in enterprise environments.
  • Experience building and operating scalable AWS-based data platforms and pipelines using services including Lambda, Glue, Athena, S3, Redshift, DMS, MWAA (Airflow), and Step Functions, supporting batch, CDC, and near real-time data processing.
  • Advanced proficiency in Python, SQL, and PySpark with hands‑on experience developing reusable ETL/ELT frameworks, data warehouses, data marts, and integrations across databases, APIs, event streams, and analytics environments.
  • Experience implementing data quality, governance, and optimization best practices, including automated validation frameworks, Lake Formation and Glue Data Catalog, performance tuning, and cost optimization across AWS data services.
  • Strong communication skills with the ability to translate complex data concepts for business stakeholders; experience in healthcare, life sciences, and other highly regulated environments with HIPAA, GDPR, FDA, or similar compliance requirements preferred.
  • Experience with metadata management, data lineage, data observability, master data management, or enterprise data catalog solutions.
  • Knowledge with data modeling & analytical data model, schema design, schema evolution, and data structure optimized for reporting and analytics.
  • Knowledge of data lake and data warehouse architecture include data partitioning and columnar storage format such as Parquet.
  • Relevant AWS certification, such as AWS Certified Data Engineer – Associate, or an equivalent cloud or data engineering certification.

WORKING CONDITIONS

  • Flexible work hours in fun collaborative environment
  • Working remote requires a reliable internet connection
  • Must have the ability to travel, as needed for company meetings
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Data Engineer
Data Engineer

The Value Maximizer • South Carolina

On-site
USD 90,000 - 120,000
Data Engineer - Python, SQL, AWS
Data Engineer - Python, SQL, AWS

Compunnel, Inc. • Durham (NC)

On-site
USD 95,000 - 120,000
Data Engineer
Data Engineer

Compunnel, Inc. • Boston (MA)

On-site
USD 110,000 - 140,000
AWS Lakehouse Data Engineer
AWS Lakehouse Data Engineer

FM Talent Source • Silver Spring (MD)

On-site
USD 140,000 - 190,000
Data Engineer
Data Engineer

Applied Resource Group • Atlanta (GA)

On-site
USD 110,000 - 130,000
Data Engineer
Data Engineer

Compunnel, Inc. • Northern (KY)

Hybrid
USD 85,000 - 125,000
Data Engineer
Data Engineer

Magpie Health Analytics, Inc. • Baltimore (MD)

On-site
USD 110,000 - 165,000
Data Engineer, AWS DC Central Operations
Data Engineer, AWS DC Central Operations

Amazon Web Services (AWS) • Seattle (WA)

On-site
USD 101,000 - 160,000
Health insurance
401(k) matching
Paid time off
+2
AWS Data Engineer
AWS Data Engineer

MMD Services, Inc • Rosemont (IL)

Hybrid
USD 135,000 - 150,000
401(k)
401(k) matching
Dental insurance
+5
AWS Data Engineer
AWS Data Engineer

Qode • United States

Remote
USD 120,000 - 150,000