Senior Developer - Python and Spark

Citi

Pune District

On-site

INR 1,800,000 - 3,000,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Hybrid work model
Growth opportunities
Global engineering network
Wellbeing program
Competitive benefits

Job summary

Citi in India is seeking a data platform engineer to design and build scalable ETL/ELT pipelines in Python and PySpark, powering reliable data ingestion and transformation. You will implement AI-first approaches, integrate ML models, and collaborate with engineers and product owners to accelerate outcomes.

The role emphasizes data quality, performance tuning, and DevOps practices. Hybrid work model offers in-person collaboration 3 days a week while enabling remote work, with opportunities to

Qualifications

  • Deep hands-on expertise in Python and Spark, including Spark Core/SQL and DataFrames.
  • Advanced SQL for complex queries and performance optimization.
  • Experience building large-scale ETL/ELT pipelines.
  • Proficiency with Python data libraries for wrangling and preprocessing.
  • DevOps mindset with CI/CD and GitHub version control.
  • Solid Linux knowledge and familiarity with Windows.
  • Ability to collaborate across multidisciplinary teams to solve problems.

Responsibilities

  • Design and build scalable ETL/ELT pipelines in Python and PySpark.
  • Develop production-grade code and mentor team through reviews.
  • Integrate Generative AI/ML models into data workflows.
  • Enforce data quality standards across pipelines.
  • Adopt AI-first engineering and Agentic AI workflows.
  • Automate CI/CD pipelines for builds and deployments.
  • Tune performance for large volumes of data.
  • Collaborate with engineers and product owners to align goals.

Skills

Python & Spark
Advanced SQL
ETL/ELT pipelines
Pandas Numpy SciPy
CI/CD GitHub
Linux
Cross-functional teamwork

Tools

GitHub
Linux
Jupyter Notebook

Job description

Responsibilities
  • Design and build highly scalable ETL/ELT pipelines in Python and PySpark to power reliable data ingestion, transformation, and integration at scale.
  • Develop production‑grade code and elevate team capability through structured code reviews and hands‑on peer mentorship.
  • Integrate Generative AI and Machine Learning models into data workflows, using AI development tools to deliver innovative solutions and resolve complex technical challenges.
  • Establish and enforce data quality standards across all pipelines, ensuring accuracy, consistency, and reliability of data throughout its lifecycle.
  • Implement an AI‑first engineering approach by exploring and embedding Agentic AI workflows to increase development velocity and platform intelligence.
  • Automate CI/CD pipelines for builds and deployments, advancing DevOps maturity across the data platform.
  • Apply performance tuning and optimisation techniques to process large volumes of structured and semi‑structured data efficiently and at speed.
  • Collaborate directly with engineers and product owners to align technical direction with delivery goals and accelerate outcomes.
Required Qualifications & Skills
  • Deep hands‑on expertise in Python and Apache Spark, including Spark Core, Spark SQL, and DataFrames/Datasets, applied to production data systems.
  • Advanced SQL capability across complex query writing, data transformation logic, and query performance optimisation.
  • Demonstrated experience designing, building, and maintaining ETL/ELT pipelines for large‑scale data ingestion and transformation.
  • Proficiency with Python data processing libraries including Pandas, NumPy, and SciPy for data wrangling and preprocessing tasks.
  • Practical experience with DevOps tools and CI/CD pipeline automation, with version control managed through GitHub.
  • Solid working knowledge of Linux environments, with familiarity across Windows systems.
  • Ability to collaborate across multidisciplinary teams to diagnose complex technical problems and deliver effective, lasting solutions.
Beneficial Skills & Qualifications
  • Working knowledge of Machine Learning and Generative AI, with direct exposure to AI development tooling in an engineering context.
  • Familiarity with Agentic AI workflows and their application to enhancing platform capabilities or development productivity.
  • Experience using Jupyter Notebook for rapid prototyping and iterative development of data solutions.
What We Offer
  • Hybrid working model with 3 days in the office and 2 days working remotely, giving you flexibility and in‑person collaboration.
  • Access to complex, large‑scale engineering challenges at the intersection of data and AI, keeping your technical skills at the forefront of the industry.
  • Opportunities to grow into AI and platform architecture disciplines, supported by a team culture that invests in continuous technical development.
  • Exposure to a globally connected engineering network, collaborating with specialists across technology, data, and product functions.
  • A performance‑driven environment where your contributions to platform innovation directly influence business outcomes at scale.
  • Wellbeing and work‑life balance support, alongside a competitive benefits package designed to recognise the value you bring.

Build data systems that power real decisions — apply now to join Citi's next‑generation AI and data platform engineering team.

Citi is an equal opportunity employer, and qualified candidates will receive consideration without regard to their race, color, religion, sex, sexual orientation, gender identity, national origin, disability, status as a protected veteran, or any other characteristic protected by law.

If you are a person with a disability and need a reasonable accommodation to use our search tools and/or apply for a career opportunity review Accessibility at Citi.

View Citi’s EEO Policy Statement and the Know Your Rights poster.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Developer - Python and Spark
Senior Developer - Python and Spark

Citigroup Inc. • Pune District

Hybrid
INR 2,800,000 - 4,200,000
Hybrid work arrangement
Exposure to AI/data platform projects
Growth in AI and platform architecture
Senior Developer - Python and Spark
Senior Developer - Python and Spark

Citi • Maharashtra

Hybrid
INR 1,200,000 - 1,800,000
Hybrid work schedule
Senior Developer - Python and Spark
Senior Developer - Python and Spark

Citibank (Switzerland) AG • Pune District

Hybrid
Confidential
Hybrid work model (3 days in office, 2
Global engineering network
Wellbeing and work-life balance
Senior Data Engineer - Vice President - Python Development
Senior Data Engineer - Vice President - Python Development

Citi • Pune District

On-site
INR 3,000,000 - 6,000,000
Python and AI Data Engineering Lead - Senior Vice President
Python and AI Data Engineering Lead - Senior Vice President

Citi • Pune District

On-site
INR 2,800,000 - 4,600,000
Python and AI Data Engineering Lead - Senior Vice President
Python and AI Data Engineering Lead - Senior Vice President

Citi • Maharashtra

On-site
INR 3,000,000 - 5,500,000
Python data engineer - AI and ML applications - Vice president
Python data engineer - AI and ML applications - Vice president

Citibank (Switzerland) AG • Chennai District

On-site
Confidential
AI Python Engineer - Vice President
AI Python Engineer - Vice President

Citigroup Inc. • Pune District

On-site
INR 1,800,000 - 2,400,000
Senior Data Engineer - Vice President - Python Development
Senior Data Engineer - Vice President - Python Development

Citigroup Inc. • Pune District

On-site
INR 3,500,000 - 9,000,000
Python and AI Data Engineering Lead - Senior Vice President
Python and AI Data Engineering Lead - Senior Vice President

Citigroup Inc. • Pune District

On-site
INR 2,500,000 - 4,500,000