AI-Driven Lead Data Engineer (AWS/EMR)

Millennium Global Technologies

Dallas (TX)

Hybrid

USD 140,000 - 210,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Millennium Global Technologies is seeking a Lead Data Engineer to own and optimize the HR data platform, building and refining Spark/ETL pipelines, semantic layers, and RAG workflows. You will partner with US HR stakeholders and offshore IDC teams to deliver reliable data assets and AI-enabled insights.

The role demands 8+ years of deep data engineering experience, strong Python/Scala, SQL, and cloud data warehousing expertise, with a proven ability to operate end-to-end and communicate across

Qualifications

  • 8+ years in Data Engineering with hands-on Spark (PySpark/Scala), SQL and Python.
  • Experience building and operating ETL/ELT pipelines on cloud platforms (AWS EMR/S3, Databricks) and workflow orchestration (Airflow).
  • Comfortable owning pipeline operations end-to-end: reading dependencies, diagnosing failures, coordinating with multiple teams.
  • Knowledge of data warehousing/data marts, governance, and PII handling practices.
  • AI-native mindset with familiarity with GenAI tools to enhance engineering workflows.
  • Strong communicator and experience directing offshore teams (IDC).

Responsibilities

  • Build, maintain, and troubleshoot Spark/EMR ETL pipelines feeding HR/workforce data marts.
  • Monitor data quality and governance rules to keep HR data assets compliant.
  • Coordinate with US HR stakeholders and IDC engineering teams; translate requirements and report status.
  • Triage Jira tickets; write runbooks and pipeline documentation.
  • Build semantic layers and Retrieval-Augmented Generation pipelines.
  • Integrate REST/GraphQL APIs into data workflows.
  • Operate with minimal oversight; flag risks proactively.

Skills

Spark
PySpark/Scala
SQL
Python
AWS
Airflow
Data Warehousing
Data Governance
PII handling
GenAI tools

Tools

AWS EMR
S3
Databricks
REST APIs
GraphQL
PagerDuty
Splunk

Job description

Millennium Global Technologies is seeking a Lead Data Engineer to own and optimize the HR data platform, building and refining Spark/ETL pipelines, semantic layers, and RAG workflows. You will partner with US HR stakeholders and offshore IDC teams to deliver reliable data assets and AI-enabled insights.

The role demands 8+ years of deep data engineering experience, strong Python/Scala, SQL, and cloud data warehousing expertise, with a proven ability to operate end-to-end and communicate across

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

AI Lead Data Engineer (No GC''s)
AI Lead Data Engineer (No GC''s)

Millennium Global Technologies • Dallas (TX)

Hybrid
USD 140,000 - 210,000
Lead Data Engineer
Lead Data Engineer

Harnham • Dallas (TX)

On-site
USD 130,000 - 160,000
AWS Data Engineer | Big Data & Cloud Data Engineer
AWS Data Engineer | Big Data & Cloud Data Engineer

Veriipro • Englewood (CO)

On-site
USD 120,000 - 170,000
Lead Data Engineer - Big Data
Lead Data Engineer - Big Data

Compunnel, Inc. • Atlanta (GA)

On-site
USD 120,000 - 150,000
Big Data Lead
Big Data Lead

Veriipro • United States

On-site
USD 180,000 - 240,000
Lead Data Engineer (AI & Data Platforms)
Lead Data Engineer (AI & Data Platforms)

Eleven Recruiting • El Segundo (CA)

On-site
USD 180,000 - 260,000
Lead Data Engineer
Lead Data Engineer

TechDigital Group • San Jose (CA)

On-site
USD 130,000 - 160,000
Lead Data Engineer
Lead Data Engineer

Recru • Houston (TX)

On-site
USD 130,000 - 150,000
Senior Data Engineer
Senior Data Engineer

TechDigital Group • Denver (CO)

On-site
USD 120,000 - 170,000
Senior Data Engineer - AI/ML & Spark Expert
Senior Data Engineer - AI/ML & Spark Expert

Genius Business Solutions • Aurora (CO)

On-site
USD 110,000 - 150,000