PySpark Big Data Developer

Citi

Maharashtra

On-site

INR 1,100,000 - 1,800,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Citi is seeking a highly skilled Big Data/PySpark Engineer to design, develop, and optimize scalable data pipelines for large-scale analytics. This role sits in a dynamic Big Data Analytics team focused on robust, scalable data processing and analytics.

Responsibilities include building PySpark ETL pipelines, writing clean Python code, collaborating with data engineers and analysts, debugging Spark jobs, and contributing to the full SDLC while ensuring data quality and integrity.

Qualifications

  • 4–8 years of experience in developing and managing enterprise data pipelines.
  • Strong knowledge of PySpark for large-scale data processing.
  • Experience with Hadoop ecosystem (HDFS, Hive, Sqoop).
  • Proficient with SQL Server and Oracle databases; ability to write data validation queries.
  • Scripting and scheduling experience with Shell scripts and Autosys.
  • Familiarity with BI tools such as Tableau.
  • Version control with Git; familiarity with Jira and Confluence.

Responsibilities

  • Design and develop scalable PySpark ETL pipelines for large datasets.
  • Write clean, documented Python code.
  • Collaborate with cross-functional teams to gather data requirements.
  • Debug and optimize Spark applications and data processing performance.
  • Participate in the full SDLC from requirements to deployment.
  • Ensure data quality and integrity across the data lifecycle.

Skills

PySpark
OOP concepts
Hadoop
Hive
Sqoop
SQL
Shell scripting
Tableau
Git

Tools

Autosys
JIRA
Confluence

Job description

Overview

We are seeking a highly skilled and experienced BigData/PySpark Engineer to join our dynamic Big Data Analytics team. This role is pivotal in designing, developing, and optimizing robust, scalable data pipelines for large-scale data processing and analytics.

Key Responsibilities
  • Design & Development: Create and optimize scalable ETL (Extraction, Transformation, Loading) pipelines using PySpark for massive datasets.
  • Coding & Engineering: Write clean, efficient, well-documented code primarily in Python (PySpark) often leveraging frameworks/tools.
  • Collaboration: Work with cross-functional teams (senior developers, data engineers, analysts, business partners) to understand data requirements and ensure seamless solution integration.
  • Troubleshooting & Optimization: Debug and resolve data processing issues and performance bottlenecks in Spark applications and other big data technologies.
  • Full SDLC Involvement: Participate in the entire software development lifecycle, from requirements analysis and design to testing, deployment, and operations.
  • Data Integrity: Ensure high data quality and integrity throughout the data lifecycle.
Qualifications & Experience
  • 4-8 years of experience in developing and managing enterprise applications.
  • Solid foundation in Object-Oriented Programming (OOP) concepts.
  • Expertise in PySpark, HDFS, Hive, Sqoop, and Hadoop for Big Data environments.
  • Good exposure to SQL Server and Oracle databases; experience with query writing for data validation and manipulation.
  • Proficient in Shell Scripting and job scheduling tools like Autosys.
  • Some exposure to BI tools such as Tableau.
  • Proficient with Git; experience with JIRA, Confluence; familiarity with DevOps and CI/CD pipelines.

Citi is an equal opportunity employer, and qualified candidates will receive consideration without regard to race, color, religion, sex, sexual orientation, gender identity, national origin, disability, status as a protected veteran, or any other characteristic protected by law. If you are a person with a disability and need a reasonable accommodation to use our search tools and/or apply for a career opportunity, view Citi’s EEO Policy Statement and the Know Your Rights poster.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

PySpark Big Data Developer
PySpark Big Data Developer

Citibank (Switzerland) AG • Pune District

Hybrid
Confidential
Senior Big Data Engineer Pyspark Databricks - Vice President
Senior Big Data Engineer Pyspark Databricks - Vice President

Citi • Maharashtra

On-site
INR 3,000,000 - 4,500,000
Senior Big Data Engineer Pyspark Databricks - Vice President
Senior Big Data Engineer Pyspark Databricks - Vice President

Citigroup Inc. • Pune District

On-site
INR 3,000,000 - 5,000,000
Senior Big Data Engineer Pyspark Databricks - Vice President
Senior Big Data Engineer Pyspark Databricks - Vice President

Citi • Pune District

On-site
INR 3,500,000 - 6,000,000
PySpark Big Data Developer
PySpark Big Data Developer

Citigroup Inc. • Pune District

On-site
INR 4,766,000 - 7,150,000
Big Data / PySpark Engineering Lead - Vice President
Big Data / PySpark Engineering Lead - Vice President

Citigroup Inc. • Pune District

On-site
INR 1,800,000 - 2,500,000
Big Data Engineer - Python and Spark
Big Data Engineer - Python and Spark

Citigroup Inc. • Pune District

On-site
INR 600,000 - 1,000,000
Applications Developer - ETL/Pyspark - Assistant Vice President
Applications Developer - ETL/Pyspark - Assistant Vice President

Citibank (Switzerland) AG • India

On-site
Confidential
PySpark Data Engineer
PySpark Data Engineer

Code1 Tech Systems • India

On-site
INR 1,200,000 - 2,400,000
Big Data Engineer - Python and Spark
Big Data Engineer - Python and Spark

Citibank (Switzerland) AG • Pune District

On-site
Confidential