Pyspark Developer (5 locations)

Tata Consultancy Services

Bengaluru

On-site

INR 1,800,000 - 3,200,000

Full time

3 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Tata Consultancy Services in Bengaluru is seeking an experienced data engineer to design and optimize PySpark-based ETL pipelines across multiple data sources.

Responsibilities include working with Hive/HDFS, Parquet/Avro/JSON, and cloud platforms (AWS/GCP/Azure) to process data batch and real-time. Strong SQL, Airflow or Oozie, and Git are essential; collaboration with data scientists and analysts is required to deliver scalable solutions.

Qualifications

  • Strong Python and PySpark development experience.
  • Hands-on with Spark SQL, RDD and DataFrame APIs.
  • Solid understanding of Hadoop ecosystem components (Hive, HDFS, YARN).
  • Experience with data formats such as Parquet, Avro and JSON.
  • Proficient SQL skills and writing efficient queries.
  • Familiarity with workflow tools (Airflow, Oozie) and Git.
  • Exposure to cloud platforms (AWS/GCP/Azure) is a plus.

Responsibilities

  • Design, develop, and maintain robust ETL/ELT pipelines using PySpark and other big data technologies. Optimize Spark jobs for performance and scalability.
  • Implement data transformations, aggregations, and joins over large datasets. Perform batch and real-time data processing tasks.
  • Collaborate with data scientists, analysts, and engineers to understand requirements and deliver quality solutions. Integrate PySpark with data warehouses (Hive, Redshift, Snowflake) and other data stores.
  • Write clean, maintainable, tested code. Follow version control and CI/CD practices using Git, Jenkins, or Azure DevOps.
  • Troubleshoot data quality and performance issues. Monitor pipeline health and ensure SLAs are met.
  • Work with platforms like Hadoop, Hive, HDFS, AWS EMR, Databricks, or Azure Synapse as required.

Skills

Python
PySpark
Spark SQL
RDD
DataFrame APIs
Hadoop ecosystem
Hive
HDFS
YARN
Parquet
Avro
JSON formats
SQL
Airflow
Oozie
Git
Cloud platforms (AWS/GCP/Azure)

Education

15 years of Full Time Education

Tools

Jenkins
Azure DevOps

Job description

Experience Range: - 05 To 10 Years (Mandatory)

(Note: Candidates below 5 years of IT experience shall not be considered)

Job Locations: Bengaluru, Chennai, Hyderabad, Pune, Kolkata

JOB DESCRIPTION
Job Requirements:
  • Experience with strong proficiency in Python and PySpark.
  • Experience working on Spark SQL, RDD, and DataFrame APIs.
  • Good understanding of Hadoop ecosystem (Hive, HDFS, YARN).
  • Knowledge of data formats like Parquet, Avro, JSON, etc.
  • Experience with SQL and writing efficient queries.
  • Familiarity with job orchestration tools (Airflow, Oozie, or similar).
  • Version control systems like Git.
  • Exposure to cloud platforms (AWS/GCP/Azure) is a plus.
Key responsibilities:
  • Design, develop, and maintain robust ETL/ELT pipelines using PySpark and other big data technologies. Optimize Spark jobs for performance and scalability.
  • Implement data transformations, aggregations, and joins over large datasets. Perform batch and real-time data processing tasks.
  • Collaborate with data scientists, analysts, and other engineers to understand requirements and deliver quality solutions. Integrate PySpark solutions with data warehouses (like Hive, Redshift, Snowflake) and other data stores.
  • Write clean, maintainable, and well-documented code. Follow version control and CI/CD practices using tools like Git, Jenkins, or Azure DevOps.
  • Troubleshoot data quality and performance issues. Monitor pipeline health and ensure SLAs are met.
  • Work with tools and platforms like Hadoop, Hive, HDFS, AWS EMR, Databricks, or Azure Synapse as required.

15 years of Full Time Education

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Pyspark (Data Engineer)-Pan India-2-5yrs
Pyspark (Data Engineer)-Pan India-2-5yrs

Infosys • Mumbai, Chennai District, Pune District

Hybrid
INR 1,200,000 - 2,200,000
Pyspark Developer
Pyspark Developer

ZettaMine Labs • Pune District

On-site
INR 1,000,000 - 1,800,000
Python & PySpark Developer ( 5-7 Years ) Hyderabad
Python & PySpark Developer ( 5-7 Years ) Hyderabad

HypTechie • Hyderabad

On-site
INR 1,500,000 - 2,100,000
Pyspark Spark Python Developer - Chennai/ Bangalore
Pyspark Spark Python Developer - Chennai/ Bangalore

Tech Mahindra • Chennai District

On-site
INR 1,500,000 - 2,500,000
Pyspark Developer-Pan India-2-5yrs
Pyspark Developer-Pan India-2-5yrs

Infosys • Chennai District, Bengaluru, Hyderabad

On-site
INR 700,000 - 1,100,000
Python + Hadoop + Pyspark
Python + Hadoop + Pyspark

NITYO • Bengaluru

On-site
INR 1,200,000 - 1,800,000
Pyspark developer
Pyspark developer

Aligned Automation • Pune District

On-site
INR 2,200,000 - 3,400,000
PySpark Developer / Senior Data Engineer
PySpark Developer / Senior Data Engineer

Alignity Solutions • Hyderabad

On-site
INR 1,200,000 - 1,800,000
Pyspark - Bigdata Engineer
Pyspark - Bigdata Engineer

Virtusa • Chennai District, Pune District

On-site
INR 1,000,000 - 1,800,000
Developer - PySpark
Developer - PySpark

Compunnel, Inc. • Pune District

On-site
INR 800,000 - 1,500,000