Aws Data Engineer

Tranzeal

Bengaluru

Hybrid

INR 1,200,000 - 1,800,000

Full time

11 days ago

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Tranzeal is seeking a skilled Data Engineer to design, build, and maintain scalable data pipelines powering analytics, reporting, and ML initiatives. You will work extensively with PySpark and the AWS ecosystem to create reliable, cloud-native data infrastructure.

Responsibilities include building ETL/ELT pipelines, managing data lakes and warehouses on AWS, orchestrating workflows with Airflow or Glue, and ensuring data quality, security, and governance across production pipelines.

Qualifications

  • Experience designing and maintaining ETL/ELT pipelines with PySpark.
  • Building data lake/warehouse solutions on AWS.
  • Orchestrating workloads with Airflow or AWS Glue Workflows.
  • Optimizing Spark jobs for performance and cost.
  • Collaborating with data analysts and data scientists.

Responsibilities

  • Ingest data from APIs, databases, flat files, and streaming platforms like Kafka/Kinesis.
  • Implement data quality checks, validation frameworks, and monitoring for pipeline health.
  • Collaborate with data analysts, data scientists, and business stakeholders to understand data requirements.
  • Design and maintain data models (star/snowflake schemas) for analytics use cases.
  • Write clean, well-documented, testable code following CI/CD and version control.
  • Ensure data security, governance, and compliance (IAM policies, encryption, access controls).
  • Troubleshoot and resolve production data pipeline issues.

Skills

PySpark
AWS
ETL/ELT
Spark tuning
Data modeling
CI/CD
Kafka/Kinesis
Data quality
Collaboration

Tools

Apache Airflow
AWS Glue
Amazon Redshift
S3
EMR
Lake Formation
Kinesis

Job description

JD:
We're looking for a skilled Data Engineer to design, build, and maintain scalable data pipelines that power analytics, reporting, and machine learning initiatives. You'll work extensively with PySpark for large-scale data processing and the AWS ecosystem to build reliable, cloud-native data infrastructure.

  • Design, develop, and maintain ETL/ELT pipelines using PySpark for batch and streaming data processing
  • Build and manage data lake and data warehouse solutions on AWS (S3, Redshift, Glue, EMR, Athena, Lake Formation)
  • Develop and orchestrate workflows using AWS Step Functions, Apache Airflow, or AWS Glue Workflows
  • Optimize Spark jobs for performance, cost, and scalability (partitioning, caching, cluster tuning)
  • Ingest data from multiple sources (APIs, databases, flat files, streaming platforms like Kafka/Kinesis)
  • Implement data quality checks, validation frameworks, and monitoring/alerting for pipeline health
  • Collaborate with data analysts, data scientists, and business stakeholders to understand data requirements
  • Design and maintain data models (star/snowflake schemas) for analytics use cases
  • Write clean, well-documented, testable code following engineering best practices (CI/CD, version control)
  • Ensure data security, governance, and compliance (IAM policies, encryption, access controls)
  • Troubleshoot and resolve production data pipeline issues
Get your free, confidential resume review.
or drag and drop your file here.