Remote AWS Lakehouse Data Engineer

Delan Associates, Inc

Atlanta (GA)

Remote

USD 120,000 - 180,000

Full time

12 days ago
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Job summary

Delan Associates, Inc. is seeking an AWS Lakehouse Data Engineer to design and operate a cloud-native data platform powering AI/ML, analytics, reporting, and data visualization.

You will build a modern lakehouse on Amazon S3 using AWS-native services and open table formats, ensuring portability, governance, cost efficiency, and operational control. This role emphasizes scalable batch and streaming ingestion, Python and PySpark ETL/ELT pipelines, metadata and governance services, and automated

Qualifications

  • Bachelor's degree in Engineering, IT, CS, Data Engineering, or 4+ years equivalent practical experience.
  • SIX (6) years of relevant experience.
  • Hands-on experience implementing AWS-native data lake or lakehouse architectures using Amazon S3 and services such as AWS Glue, Amazon Athena, Amazon EMR, AWS Lake Formation, and Amazon Redshift.
  • Strong experience developing production ETL/ELT pipelines using Python and PySpark, including data modeling, transformation, testing, performance tuning, and error handling.
  • Hands-on experience with Apache Iceberg, including ACID transactions, snapshots, schema and partition evolution, time travel, table maintenance, and query optimization.
  • Advanced SQL skills and experience supporting analytical queries, semantic layers, reporting tools, and data visualization workloads.
  • Experience implementing metadata management and governance capabilities, including cataloging, lineage, ownership, classification, policy enforcement, and fine-grained access controls.
  • Experience with AWS security fundamentals, including IAM and least privilege, KMS encryption, secrets management, network security, logging, and secure SDLC practices.
  • Experience provisioning AWS resources using IaC and operating data platforms across multiple environments.
  • Experience building or operating CI/CD pipelines for data workflows, including testing, packaging, deployment automation, environment promotion, and rollback.

Responsibilities

  • Build and Operate Data Pipelines (Batch and Streaming). Design and implement batch and streaming ingestion from APIs, relational databases, file drops, event streams, and external partners.
  • Implement, test, and optimize ETL/ELT pipelines using Python and PySpark to produce curated, analytics-ready datasets for reporting, visualization, and machine learning.
  • Implement incremental processing, change data capture (CDC), data contracts, schema validation, and reusable transformation frameworks.
  • Improve pipeline reliability through automated testing, orchestration, monitoring, retry handling, and operational runbooks.
  • Deliver an AWS-Native Lakehouse Data Platform.
  • Design and implement a Delta Lakehouse-style data platform using AWS-native services to provide Databricks-like capabilities for data engineering, analysis, and data visualization.
  • Build and manage a scalable lakehouse on Amazon S3 using Apache Iceberg and open columnar formats such as Apache Parquet.
  • Implement SQL-like table reliability for data stored in Amazon S3, including ACID transactions, schema evolution, partition evolution, snapshot isolation, time travel, and rollback capabilities using Apache Iceberg.
  • Enable fast, interactive querying of lakehouse data using AWS-native query and compute services such as Amazon Athena, Amazon EMR, AWS Glue, and Amazon Redshift where appropriate.
  • Optimize performance and cost through partitioning, compaction, file sizing, statistics, caching, lifecycle policies, and efficient separation of compute and storage.
  • Establish standardized development, test, and production environments with consistent configuration and controlled promotion across stages.
  • Metadata, Governance, Access Control, Lineage, and Quality.
  • Implement data governance and fine-grained access control using AWS-native services, including AWS Lake Formation, AWS Glue Data Catalog, AWS Identity and Access Management (IAM), AWS Key Management Service (KMS), and related security services.
  • Implement a managed metadata repository for dataset cataloging, ownership, business definitions, tagging, classification, and discoverability.
  • Enable end-to-end lineage from source through transformation and consumption to support auditability, impact analysis, and regulatory requirements.
  • Apply policy-based access, least-privilege permissions, row-, column-, and cell-level controls where required, data classification, retention, encryption, and secure data handling.
  • Build operational data quality checks for freshness, completeness, uniqueness, validity, consistency, and anomaly detection, and publish measurable SLAs/SLOs.
  • AWS Automation, CI/CD, and Operations.
  • Implement automated AWS provisioning using Infrastructure as Code (IaC) to create consistent environments and secure-by-default baselines.
  • Build and enhance CI/CD for data pipelines and lakehouse components, including automated tests, security checks, validation gates, packaging, deployment, promotion, and rollback strategies.
  • Implement observability with centralized metrics, logs, traces, alerts, dashboards, runbooks, and incident-response procedures.
  • Continuously evaluate platform performance, scalability, reliability, security, and cost, and implement measurable improvements.
  • Cross-Team Collaboration and Documentation.
  • Work closely with data, application, analytics, AI/ML, security, networking, and cloud platform teams to support mission needs and delivery timelines.
  • Maintain high-quality engineering documentation, including architecture diagrams, data models, SOPs, interface specifications, operational runbooks, and secure configuration baselines.
  • Present technical findings, trade-offs, risks, and recommendations clearly to technical and non-technical stakeholders.

Skills

Python
PySpark
SQL
ETL/ELT
Apache Iceberg
AWS
Data governance
IAM/KMS security
IaC
CI/CD
Spark

Education

Bachelor's degree in a related field

Tools

Databricks
Delta Lake
Terraform
CloudFormation
CDK
Jenkins
GitHub Actions
Docker

Job description

Delan Associates, Inc. is seeking an AWS Lakehouse Data Engineer to design and operate a cloud-native data platform powering AI/ML, analytics, reporting, and data visualization.

You will build a modern lakehouse on Amazon S3 using AWS-native services and open table formats, ensuring portability, governance, cost efficiency, and operational control. This role emphasizes scalable batch and streaming ingestion, Python and PySpark ETL/ELT pipelines, metadata and governance services, and automated

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Remote AWS Lakehouse Data Engineer for AI/ML Pipelines
Remote AWS Lakehouse Data Engineer for AI/ML Pipelines

Delan Associates, Inc • United States

Remote
USD 140,000 - 180,000
Remote AWS Lakehouse Data Engineer—ETL, Governance
Remote AWS Lakehouse Data Engineer—ETL, Governance

Delan Associates, Inc. • Atlanta (GA)

Remote
USD 140,000 - 200,000
AWS Lakehouse Data Engineer for AI/ML Analytics
AWS Lakehouse Data Engineer for AI/ML Analytics

Dovel Technologies, Inc • Northern (KY)

Hybrid
USD 113,000 - 188,000
Medical Insurance
401(k) Retirement Plan
Tuition Reimbursement
+2
AWS Lakehouse Data Engineer: Scalable AI/ML Pipelines
AWS Lakehouse Data Engineer: Scalable AI/ML Pipelines

Guidehouse • Washington

On-site
USD 113,000 - 188,000
Medical Insurance
401(k) Retirement Plan
Paid Holidays
Senior AWS Data Engineer: Lakehouse & AI-Driven Pipelines
Senior AWS Data Engineer: Lakehouse & AI-Driven Pipelines

Global Business Ser. 4u • New York (NY)

On-site
USD 140,000 - 190,000
AWS Lakehouse Data Engineer- GrantSolutions experienced only
AWS Lakehouse Data Engineer- GrantSolutions experienced only

TRIWAVE Solutions INC • United States

Remote
USD 80,000 - 113,000
AWS Lakehouse Data Engineer: Build Scalable Platform
AWS Lakehouse Data Engineer: Build Scalable Platform

Guidehouse • United States

Remote
USD 113,000 - 188,000
Medical, Rx, Dental & Vision Insurance
401(k) Retirement Plan
Parental Leave
+4
Databricks Lakehouse Engineer: PySpark, Delta, AI/ML
Databricks Lakehouse Engineer: PySpark, Delta, AI/ML

Delan Associates, Inc • Louisville (KY)

On-site
USD 100,000 - 150,000
Remote AWS Lakehouse Data Engineer | Pipelines & Governance
Remote AWS Lakehouse Data Engineer | Pipelines & Governance

TRIWAVE Solutions INC • United States

Remote
USD 80,000 - 113,000
Data Engineer – AWS Lakehouse (Mandarin Required)
Data Engineer – AWS Lakehouse (Mandarin Required)

Bitus Labs • Irvine (CA)

On-site
USD 120,000 - 180,000