Data Engineer

KKR

Gurugram District

On-site

INR 1,800,000 - 2,800,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

KKR is seeking an experienced Data Engineer (5‑8 years) to design, build, and operate scalable data platforms on AWS. You will own the full lifecycle of data pipelines—from ingestion to governance—using a modern lakehouse stack built on Apache Iceberg, AWS Glue, and Snowflake.

You will partner with data scientists, analysts, and platform teams to deliver reliable, well‑governed data products that power analytics and business decisions across Insurance Systems.

Qualifications

  • 5-8 years of hands‑on data engineering experience building production data pipelines.
  • Strong proficiency in Python for data engineering and automation.
  • Advanced SQL and strong experience with relational databases (e.g., PostgreSQL, MySQL).
  • A data pipeline orchestrator is mandatory - experience with a workflow orchestration tool in production.
  • Hands‑on experience with Apache Spark for large‑scale distributed data processing.
  • Production experience with Apache Iceberg for lakehouse storage.
  • Hands‑on experience with AWS Glue (ETL jobs and the Glue Data Catalog).
  • Experience with Snowflake as a cloud data warehouse.
  • Experience with data cataloging and governance using AWS Glue Data Catalog and Snowflake Horizon.
  • Strong experience with AWS as the primary cloud platform (S3, EMR, Lambda, Athena, Kinesis, Redshift, IAM).
  • Knowledge of data governance, data warehousing, and lakehouse patterns.
  • Experience with REST APIs and data integration techniques.

Responsibilities

  • Design, build, and maintain scalable, fault‑tolerant data pipelines and ETL/ELT processes.
  • Own end‑to‑end pipeline orchestration including scheduling, retries, SLAs, and observability.
  • Build and manage lakehouse data assets on Iceberg, including partitioning and schema evolution.
  • Develop and operate large‑scale data processing jobs using Apache Spark and AWS Glue Spark jobs.
  • Model, load, and optimize data in Snowflake, and govern data assets using data catalogs.
  • Architect and implement data solutions on AWS using S3, Glue, EMR, Lambda, Athena, Kinesis, Redshift, IAM.
  • Ensure data quality, lineage, and security across all systems and environments.
  • Integrate data from diverse sources including APIs, databases, and streaming feeds.
  • Monitor, troubleshoot, and tune pipelines for performance and cost efficiency.
  • Contribute to platform architecture, code reviews, and engineering best practices.
  • Mentor junior engineers and drive standards across the data engineering function.

Skills

Python
SQL
Airflow
Dagster
Apache Spark
AWS Glue
Iceberg
Snowflake
REST APIs
Data Modeling
ETL/ELT
Big Data

Tools

Apache Airflow
Dagster
Docker
Kubernetes
Terraform
CloudFormation/CDK

Job description

KKR's Technology team is responsible for building and supporting the firm's technological foundation including a globally distributed infrastructure, information security, and application and data platforms. The team drives a culture of technology excellence across the firm through efficient workflow automation, democratization of data through modern data and collaboration platforms, and more recently through research and development of Generative AI based tools and services.

Technology is regarded as a key business enabler at KKR and is an important accelerator to drive towards global scale creation and business process transformation. A dedicated Program Management function along with the Product Managers drive execution discipline across multiple technology teams with a goal to consistently deliver excellence serving our business needs. The Technology team consists of highly technical and business centric technologists with the ability to form strong partnerships across all of our businesses.

POSITION SUMMARY

We are seeking an experienced Data Engineer (5-8 years) to design, build, and operate scalable, production-grade data platforms on AWS. In this role you will own the full lifecycle of data pipelines - from ingestion and orchestration through transformation, storage, and governance - using a modern lakehouse stack built on Apache Iceberg, AWS Glue, and Snowflake. You will partner with data scientists, analysts, and platform teams to deliver reliable, well-governed, and cost-efficient data products that power analytics and business decisions across Insurance Systems.

ROLES & RESPONSIBILITIES
  • Design, build, and maintain scalable, fault-tolerant data pipelines and ETL/ELT processes for structured, semi-structured, and unstructured data.
  • Own end-to-end pipeline orchestration - scheduling, dependency management, retries, SLAs, and observability - using a data pipeline orchestrator.
  • Build and manage lakehouse data assets on Apache Iceberg, including partitioning, schema evolution, time-travel, and table maintenance (compaction, snapshot expiry).
  • Develop and operate large-scale data processing jobs using Apache Spark, including AWS Glue Spark jobs.
  • Model, load, and optimize data in Snowflake, and govern data assets using data catalogs (AWS Glue Data Catalog and Snowflake Horizon).
  • Architect and implement data solutions on AWS using native services (S3, Glue, EMR, Lambda, Athena, Kinesis, Redshift, IAM).
  • Ensure data quality, integrity, lineage, and security across all systems and environments.
  • Integrate data from diverse sources including APIs, relational databases, streaming feeds, and third‑party tools.
  • Monitor, troubleshoot, and tune pipelines for performance, reliability, and cost efficiency.
  • Contribute to platform architecture, design reviews, code reviews, and engineering best practices.
  • Mentor junior engineers and help drive standards across the data engineering function.
QUALIFICATIONS
  • 5-8 years of hands‑on data engineering experience building production data pipelines.
  • Strong proficiency in Python for data engineering and automation.
  • Advanced SQL and strong experience with relational databases (e.g., PostgreSQL, MySQL).
  • A data pipeline orchestrator is mandatory - proven experience operating a workflow orchestration tool (e.g., Apache Airflow, Dagster, or equivalent) in production.
  • Hands‑on experience with Apache Spark for large‑scale distributed data processing.
  • Production experience with Apache Iceberg (or an equivalent open table format) for lakehouse storage.
  • Hands‑on experience with AWS Glue (ETL jobs and the Glue Data Catalog).
  • Experience with Snowflake as a cloud data warehouse.
  • Experience with data cataloging and governance using AWS Glue Data Catalog and Snowflake Horizon.
  • Strong experience with AWS as the primary cloud platform, including S3, EMR, Lambda, Athena, Kinesis, Redshift, and IAM.
  • Solid understanding of data modeling, data warehousing, and lakehouse architecture patterns.
  • Working knowledge of REST APIs and data integration techniques.
  • Strong problem‑solving, analytical, and debugging skills.
Preferred / Good-to-Have
  • Experience with Dagster as a data pipeline orchestrator (strongly preferred).
  • Experience with containerization and orchestration (Docker, Kubernetes / Amazon EKS).
  • Exposure to CI/CD pipelines and infrastructure-as-code (e.g., Terraform, AWS CloudFormation/CDK).
  • Experience with streaming / real‑time data (Kafka, Amazon Kinesis, Spark Structured Streaming).
  • Familiarity with data observability and quality frameworks (e.g., Great Expectations, dbt tests).
  • Knowledge of data governance, security, and compliance best practices.
  • Experience within financial services or insurance data domains.
Technology Stack at a Glance
  • Cloud Platform - AWS (S3, Glue, EMR, Lambda, Athena, Kinesis, Redshift, IAM)
  • Orchestration - Data pipeline orchestrator required (e.g., Airflow); Dagster good-to-have
  • Lakehouse / Storage - Apache Iceberg on Amazon S3
  • Data Warehouse - Snowflake
  • Data Catalog / Governance - AWS Glue Data Catalog, Snowflake Horizon
Soft Skills
  • Strong communication and cross-functional collaboration skills.
  • Ability to work in a fast-paced, agile environment.
  • Self-driven with a proactive, ownership-oriented mindset.
  • Ability to mentor peers and communicate technical concepts to non-technical stakeholders.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Data Engineer (SQL, PySpark, Snowflake, DBT, Redshift) - Associate
Data Engineer (SQL, PySpark, Snowflake, DBT, Redshift) - Associate

KKR • Gurugram District

On-site
INR 2,000,000 - 3,200,000
Data Engineer (SQL, PySpark, Snowflake, DBT, Redshift) - AVP/VP
Data Engineer (SQL, PySpark, Snowflake, DBT, Redshift) - AVP/VP

KKR • Gurugram District

On-site
INR 3,500,000 - 6,000,000
Insurance Data Engineering - Data Engineer
Insurance Data Engineering - Data Engineer

KKR • Gurugram District

On-site
INR 1,400,000 - 2,500,000
Manager-Data Engineering-Big Data Engineering
Manager-Data Engineering-Big Data Engineering

EXL • Maharashtra

On-site
INR 4,000,000 - 7,000,000
Senior Data Engineer
Senior Data Engineer

DATAECONOMY Inc • Hyderabad

On-site
INR 1,500,000 - 2,100,000
Senior Data Engineer
Senior Data Engineer

SourcingXPress • Hyderabad

On-site
INR 1,500,000 - 2,500,000
Data Engineering Pipeline Engineer – Role Description
Data Engineering Pipeline Engineer – Role Description

Innoventes Technologies • Bengaluru

On-site
INR 1,000,000 - 1,500,000
Lead Data Engineer
Lead Data Engineer

CSM Technologies • Khordha

On-site
INR 1,200,000 - 1,800,000
Lead Data Architect (AWS)
Lead Data Architect (AWS)

ANRGI TECH Pvt. Ltd. • Bengaluru Urban

On-site
INR 4,000,000 - 7,000,000
Data Engineer (AWS)
Data Engineer (AWS)

Weekday (YC W21) • Gurugram District

On-site
INR 1,500,000 - 2,500,000