Senior Data Engineering Lead

UnitedHealth Group

Pune District

On-site

Confidential

Full time

2 days ago
Be an early applicant
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Job summary

UnitedHealth Group is seeking a senior data engineer to build reliable batch and streaming pipelines on AWS and Databricks. You will develop reusable ingestion frameworks for REST, FHIR HL7 interfaces, and various data formats, with Spark/PySpark, Delta Lake, and Iceberg expertise to support healthcare data models.

The role requires 5+ years of production data-engineering experience, strong Python/SQL, and a solid understanding of distributed processing, with privacy controls for PHI and

Qualifications

  • Graduate degree or equivalent experience.
  • 5+ years of production data-engineering experience.
  • Hands-on AWS experience with S3, IAM, VPC, ECS/EKS or Lambda, Step Functions, EventBridge, CloudWatch, Secrets Manager, and KMS.
  • Experience with Git, pull requests, CI/CD, Docker, Terraform/CloudFormation, automated testing, observability, and incident response.
  • Experience processing semi-structured data and building resilient ingestion pipelines with quality controls, error handling, and replay capability.
  • Production Databricks experience with Auto Loader, Delta Lake, Workflows, Unity Catalog, notebooks/jobs, SQL Warehouses, and cluster/job optimization.
  • Advanced Python, SQL, and Apache Spark/PySpark; strong understanding of distributed processing and performance tuning.
  • Practical Apache Iceberg knowledge, including tables, catalogs, snapshots, schema/partition evolution, compaction, and interoperability with query engines.
  • Healthcare data knowledge: FHIR R4 and/or HL7 v2, OMOP CDM, clinical terminologies, PHI, HIPAA-aligned engineering controls, and de-identification concepts.
  • Demonstrated responsible use of AI coding tools and the ability to critically review, test, and productionize generated code.

Responsibilities

  • Build reliable batch, micro-batch, and event-driven pipelines on AWS and Databricks.
  • Develop reusable ingestion frameworks for REST APIs, FHIR Bulk Export, HL7 interfaces, databases, SFTP/file exchange, JSON/NDJSON, CSV, XML, PDFs, and clinical text.
  • Implement scalable Spark/PySpark and SQL transformations, including schema inference/evolution, checkpointing, idempotency, retries, backfills, and replay.
  • Design and maintain Delta Lake and Apache Iceberg tables, including physical design, partitioning/clustering, compaction, file sizing, incremental reads/writes, performance tuning, and cost optimization.
  • Build curated healthcare data models and transformations using FHIR, HL7, OMOP CDM, and clinical terminology mappings.
  • Implement data-quality frameworks: schema validation, referential-integrity checks, business-rule testing, anomaly detection, reconciliation, completeness checks, and data-quality observability.
  • Build metadata and lineage capture across source systems, pipeline runs, code versions, transformation rules, mappings, and published data products.
  • Implement privacy-aware data processing for PHI, including access controls, masking, tokenization/pseudonymization, de-identification, and auditable handling patterns.
  • Deliver CI/CD pipelines, automated unit/integration/data tests, Terraform or CloudFormation, containerized services, monitoring, alerting, runbooks, and incident-recovery procedures.
  • Integrate and optimize governed data consumption through Databricks SQL, Athena, Snowflake, Trino, or equivalent engines.
  • Comply with all applicable Company policies, procedures, and business directives, changes including those relating to work location, team assignments, work schedules, and flexible work arrangements.

Skills

AWS
Databricks
Python
SQL
Spark/PySpark
Delta Lake
Iceberg
FHIR HL7 OMOP
CI/CD
Terraform/CloudFormation
GaP privacy controls

Education

Graduate degree

Tools

Docker
Git

Job description

Optum is a global organization that delivers care, aided by technology to help millions of people live healthier lives. The work you do with our team will directly improve health outcomes by connecting people with the care, pharmacy benefits, data and resources they need to feel their best. Here, you will find a culture guided by inclusion, talented peers, comprehensive benefits and career development opportunities. Come make an impact on the communities we serve as you help us advance health optimization on a global scale. Join us to start Caring. Connecting. Growing together.

Primary Responsibilities:
  • Build reliable batch, micro-batch, and event-driven pipelines on AWS and Databricks
  • Develop reusable ingestion frameworks for REST APIs, FHIR Bulk Export, HL7 interfaces, databases, SFTP/file exchange, JSON/NDJSON, CSV, XML, PDFs, and clinical text
  • Implement scalable Spark/PySpark and SQL transformations, including schema inference/evolution, checkpointing, idempotency, retries, backfills, and replay
  • Design and maintain Delta Lake and Apache Iceberg tables, including physical design, partitioning/clustering, compaction, file sizing, incremental reads/writes, performance tuning, and cost optimization
  • Build curated healthcare data models and transformations using FHIR, HL7, OMOP CDM, and clinical terminology mappings
  • Implement data-quality frameworks: schema validation, referential-integrity checks, business-rule testing, anomaly detection, reconciliation, completeness checks, and data-quality observability
  • Build metadata and lineage capture across source systems, pipeline runs, code versions, transformation rules, mappings, and published data products
  • Implement privacy-aware data processing for PHI, including access controls, masking, tokenization/pseudonymization, de-identification, and auditable handling patterns
  • Deliver CI/CD pipelines, automated unit/integration/data tests, Terraform or CloudFormation, containerized services, monitoring, alerting, runbooks, and incident-recovery procedures
  • Integrate and optimize governed data consumption through Databricks SQL, Athena, Snowflake, Trino, or equivalent engines
  • Comply with all applicable Company policies, procedures, and business directives, changes including those relating to work location, team assignments, work schedules, and flexible work arrangements
Eligibility
  • The candidate should have completed 12 months in the current role
  • The candidate should not be on any active CAP/ PIP
  • The performance review of the candidate must be ME & Above in the last common review
Required Qualifications:
  • Graduate degree or equivalent experience
  • 5+ years of production data-engineering experience
  • Hands-on AWS experience with S3, IAM, VPC, ECS/EKS or Lambda, Step Functions, EventBridge, CloudWatch, Secrets Manager, and KMS
  • Experience with Git, pull requests, CI/CD, Docker, Terraform/CloudFormation, automated testing, observability, and incident response
  • Experience processing semi-structured data and building resilient ingestion pipelines with quality controls, error handling, and replay capability
  • Production Databricks experience with Auto Loader, Delta Lake, Workflows, Unity Catalog, notebooks/jobs, SQL Warehouses, and cluster/job optimization
  • Advanced Python, SQL, and Apache Spark/PySpark; strong understanding of distributed processing and performance tuning
  • Practical Apache Iceberg knowledge, including tables, catalogs, snapshots, schema/partition evolution, compaction, and interoperability with query engines
  • Healthcare data knowledge: FHIR R4 and/or HL7 v2, OMOP CDM, clinical terminologies, PHI, HIPAA-aligned engineering controls, and de-identification concepts
  • Demonstrated responsible use of AI coding tools and the ability to critically review, test, and productionize generated code
Required Qualifications:
  • Healthcare or regulated-data experience

At UnitedHealth Group, our mission is to help people live healthier lives and make the health system work better for everyone. We believe everyone-of every race, gender, sexuality, age, location and income-deserves the opportunity to live their healthiest life. Today, however, there are still far too many barriers to good health which are disproportionately experienced by people of color, historically marginalized groups and those with lower incomes. We are committed to mitigating our impact on the environment and enabling and delivering equitable care that addresses health disparities and improves health outcomes - an enterprise priority reflected in our mission.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Principal Data Engineer
Principal Data Engineer

Optum • Chennai District

On-site
INR 4,000,000 - 7,000,000
Data Engineering Manager
Data Engineering Manager

Optum India • Bengaluru

On-site
INR 3,000,000 - 6,000,000
Senior Data Engineering Consultant
Senior Data Engineering Consultant

Optum India • Chennai District

On-site
INR 3,000,000 - 6,000,000
Senior Data Engineering Lead
Senior Data Engineering Lead

UnitedHealth Group • Hyderabad

On-site
Confidential
Data Engineering Lead
Data Engineering Lead

Optum India • Hyderabad

On-site
INR 1,800,000 - 3,600,000
Principal Data Engineer
Principal Data Engineer

Optum India • Chennai District

On-site
INR 1,500,000 - 1,900,000
Senior Data Engineering Analyst
Senior Data Engineering Analyst

Optum India • Hyderabad

On-site
INR 1,200,000 - 1,800,000
Data Engineer - Cloud ETL, Azure, Data Bricks, Snowflake, Pyspark
Data Engineer - Cloud ETL, Azure, Data Bricks, Snowflake, Pyspark

UnitedHealth Group • Hyderabad

On-site
Confidential
Senior Data Engineering Lead
Senior Data Engineering Lead

Optum India • Hyderabad

On-site
INR 4,500,000 - 6,500,000
Senior Data Engineer - Databricks & Snowflake
Senior Data Engineer - Databricks & Snowflake

Optum India • Hyderabad

On-site
INR 1,500,000 - 2,800,000