Pyspark Data Engineer

Tata Consultancy Services

Bengaluru

On-site

INR 1,200,000 - 1,800,000

Full time

10 days ago
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Job summary

Tata Consultancy Services is seeking a PySpark Data Engineer to design and implement scalable PySpark-based test architectures for ETL pipelines. The role involves building end-to-end data validation in Hadoop environments, leading system design for Hadoop/Hive test environments, and contributing to API-driven automation.

You'll mentor juniors, configure Spark for memory and core allocations, and work with HDFS/Hive to manage data, implement partitions, and optimize performance.

Qualifications

  • Design scalable PySpark-based test architectures for ETL/data pipelines.
  • Architect end-to-end data validation in Hadoop environments for lineage and schema evolution.
  • Lead system design for Hadoop/Hive test environments with YARN resource management.
  • Exposure to Zephyr-Jira-ServiceNow integrated test management systems and API-driven automation.
  • Design CI/CD test pipelines for PySpark/Hadoop jobs with artifact management.
  • Create data quality system designs using PySpark with Hive metadata services.
  • Design testing platforms and test data generators.
  • Mentor juniors on PySpark testing basics and contribute to strategy discussions.
  • Configure Spark sessions for memory and core allocations in local and cluster modes.
  • Handle data with HDFS and write back to Hive tables.
  • Implement partitioning and caching techniques to improve pipelines.
  • Perform performance tuning like salting and minimizing shuffling.

Responsibilities

  • Design and implement PySpark-based test architectures for ETL pipelines.
  • Architect data validation and lineage tracking in Hadoop/Hive environments.
  • Lead testing strategy and system design for distributed data platforms.
  • Collaborate with QA and development teams on automation and CI/CD.
  • Mentor teammates and drive best practices in PySpark testing.

Tools

PySpark
Hadoop
Hive
YARN
Jira
Zephyr
ServiceNow

Job description

Dear Professionals

Greetings from Tata consultancy Services,

Job Title Pyspark Data Engineer

Experiernce: 6-10 Years

Location: Chennai / Kolkata / Hyderabad / Pune

Mode of Work : Work from Office

Job description

  • Design scalable PySpark-based test architectures for ETL/data pipelines, including modular frameworks for batch processing.
  • Architect end-to-end data validation systems in Hadoop environment for lineage, schema evolution
  • Lead system design for Hadoop/Hive test environments, including YARN resource management, dynamic partitioning.
  • Exposure to Zephyr-Jira-ServiceNow integrated test management systems with experience on API-driven automation.
  • Design CI/CD test pipelines for PySpark/Hadoop jobs, incorporating artifact management, parallel execution, and blue-green deployments.
  • Create data quality system designs using PySpark integrated with Hive metadata services.
  • Design testing platforms, test data generators
  • Mentor juniors on PySpark testing basics, contribute to testing strategy discussions
  • Spark session configurations for memory and core allocations for both local and cluster manager settings
  • Data handling with distributed file systems like HDFS and writing back to hive tables
  • Implementation of Partitioning, caching techniques in organizing code for transformation pipelines
  • Performance tuning implementation like salting, minimizing shuffling
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Pyspark Developer
Pyspark Developer

Tata Consultancy Services • Kolkata District, Hyderabad, Chennai District

On-site
INR 2,600,000 - 3,800,000
Pyspark developer
Pyspark developer

Aligned Automation • Pune District

On-site
INR 2,200,000 - 3,400,000
Hadoop + Pyspark
Hadoop + Pyspark

Alike Thoughts • Hyderabad, Chennai District

On-site
INR 1,400,000 - 2,200,000
Pyspark Developer (5 locations)
Pyspark Developer (5 locations)

Tata Consultancy Services • Hyderabad, Chennai District, Bengaluru

On-site
INR 1,200,000 - 1,800,000
Pyspark Developer
Pyspark Developer

Tata Consultancy Services • Pune City

On-site
INR 800,000 - 1,500,000
Pyspark Developer
Pyspark Developer

ZettaMine Labs • Pune District

On-site
INR 1,000,000 - 1,800,000
Pyspark Developer
Pyspark Developer

Leading Global Technology Services Company • Chennai District, Bengaluru

Hybrid
INR 1,200,000 - 1,800,000
Senior Data Engineer(Hadoop + PySpark)
Senior Data Engineer(Hadoop + PySpark)

Alike Thoughts • Hyderabad, Chennai District

Hybrid
INR 2,500,000 - 4,200,000
PySpark Developer / Senior Data Engineer-MNC Client For 6 To 15Yrs imm
PySpark Developer / Senior Data Engineer-MNC Client For 6 To 15Yrs imm

Shell Infotech • Chennai District, Bengaluru, Hyderabad

Hybrid
INR 1,200,000 - 2,000,000
Pyspark Developer ( Pune /bangalore/Kolkata)
Pyspark Developer ( Pune /bangalore/Kolkata)

PwC India • Kolkata District, Pune District, Bengaluru

Hybrid
INR 1,400,000 - 1,800,000