AWS Lakehouse Data Engineer: Build a Modern Data Platform

FM Talent Source

Silver Spring (MD)

On-site

USD 140,000 - 190,000

Full time

2 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

FM Talent Source is seeking an AWS Lakehouse Data Engineer to design, implement, and operate a cloud-native data platform on AWS S3 using Iceberg and open formats. You will build batch and streaming pipelines, govern data, and enable analytics and AI/ML workloads.

The role requires 6+ years of experience, strong Python/PySpark, and hands-on with AWS-native services like Glue, Athena, EMR, Lake Formation, and Redshift. You will drive CI/CD, IaC, and secure data practices across environments.

Qualifications

  • Bachelor's degree in Engineering, Information Technology, Computer Science, Data Engineering, or a related field, or FOUR (4) years equivalent practical experience in lieu of degree.
  • SIX (6) years of relevant experience.
  • Hands-on experience implementing AWS-native data lake or lakehouse architectures using Amazon S3 and services such as AWS Glue, Amazon Athena, Amazon EMR, AWS Lake Formation, and Amazon Redshift.
  • Strong experience developing production ETL/ELT pipelines using Python and PySpark, including data modeling, transformation, testing, performance tuning, and error handling.
  • Hands-on experience with Apache Iceberg, including ACID transactions, snapshots, schema and partition evolution, time travel, table maintenance, and query optimization.
  • Advanced SQL skills and experience supporting analytical queries, semantic layers, reporting tools, and data visualization workloads.
  • Experience implementing metadata management and governance capabilities, including cataloging, lineage, ownership, classification, policy enforcement, and fine-grained access controls.
  • Experience with AWS security fundamentals, including IAM and least privilege, KMS encryption, secrets management, network security, logging, and secure SDLC practices.
  • Experience provisioning AWS resources using IaC and operating data platforms across multiple environments.
  • Experience building or operating CI/CD pipelines for data workflows, including testing, packaging, deployment automation, environment promotion, and rollback.
  • Ability to troubleshoot distributed data-processing workloads and optimize performance, reliability, and cost.

Responsibilities

  • Design and implement batch and streaming ingestion from APIs, relational databases, file drops, event streams, and external partners.
  • Implement, test, and optimize ETL/ELT pipelines using Python and PySpark to produce curated, analytics-ready datasets for reporting, visualization, and machine learning.
  • Implement incremental processing, change data capture (CDC), data contracts, schema validation, and reusable transformation frameworks.
  • Improve pipeline reliability through automated testing, orchestration, monitoring, retry handling, and operational runbooks.
  • Design and implement a Delta Lakehouse-style data platform using AWS-native services to provide Databricks-like capabilities for data engineering, analysis, and data visualization.
  • Build and manage a scalable lakehouse on Amazon S3 using Apache Iceberg and open columnar formats such as Apache Parquet.
  • Implement SQL-like table reliability for data stored in Amazon S3, including ACID transactions, schema evolution, partition evolution, snapshot isolation, time travel, and rollback capabilities using Apache Iceberg.
  • Enable fast, interactive querying of lakehouse data using AWS-native query and compute services such as Amazon Athena, Amazon EMR, AWS Glue, and Amazon Redshift where appropriate.
  • Optimize performance and cost through partitioning, compaction, file sizing, statistics, caching, lifecycle policies, and efficient separation of compute and storage.
  • Establish standardized development, test, and production environments with consistent configuration and controlled promotion across stages.

Skills

Python
PySpark
SQL
AWS
Iceberg
Data governance
IAM
KMS
CI/CD
IaC

Education

Bachelor's degree or 4 years experience

Tools

Databricks
Airflow
Jenkins
Terraform
CloudFormation
Git

Job description

FM Talent Source is seeking an AWS Lakehouse Data Engineer to design, implement, and operate a cloud-native data platform on AWS S3 using Iceberg and open formats. You will build batch and streaming pipelines, govern data, and enable analytics and AI/ML workloads.

The role requires 6+ years of experience, strong Python/PySpark, and hands-on with AWS-native services like Glue, Athena, EMR, Lake Formation, and Redshift. You will drive CI/CD, IaC, and secure data practices across environments.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

AWS Lakehouse Data Engineer
AWS Lakehouse Data Engineer

FM Talent Source • Silver Spring (MD)

On-site
USD 140,000 - 190,000
AWS Lakehouse Data Engineer for AI/ML Analytics
AWS Lakehouse Data Engineer for AI/ML Analytics

Dovel Technologies, Inc • Northern (KY)

Hybrid
USD 113,000 - 188,000
Medical Insurance
401(k) Retirement Plan
Tuition Reimbursement
+2
AWS Lakehouse Data Engineer: Scalable AI/ML Pipelines
AWS Lakehouse Data Engineer: Scalable AI/ML Pipelines

Guidehouse • Washington

On-site
USD 113,000 - 188,000
Medical Insurance
401(k) Retirement Plan
Paid Holidays
Data Engineer – AWS Lakehouse (Mandarin Required)
Data Engineer – AWS Lakehouse (Mandarin Required)

Bitus Labs • Irvine (CA)

On-site
USD 120,000 - 180,000
Lead Data Engineer
Lead Data Engineer

Harnham • Dallas (TX)

On-site
USD 130,000 - 160,000
Senior Data Engineer — AWS Lakehouse & Pipelines
Senior Data Engineer — AWS Lakehouse & Pipelines

RXinsider LTD. • Rockville (MD)

On-site
USD 112,000 - 147,000
Staff Lakehouse Platform Engineer
Staff Lakehouse Platform Engineer

Affirm • Boston (MA)

On-site
USD 204,000 - 264,000
Health coverage for you and dependents
Tech stipends
Generous vacation
Data Engineer
Data Engineer

The Value Maximizer • South Carolina

On-site
USD 90,000 - 120,000
Staff Lakehouse Platform Engineer
Staff Lakehouse Platform Engineer

Affirm • Chicago (IL)

On-site
USD 204,000 - 264,000
Health care coverage
Flexible Spending Wallets
Time off
+1
AWS Data Engineer
AWS Data Engineer

Capgemini • Newark (NJ)

On-site
USD 120,000 - 160,000