Senior Databricks Engineer, Apache Spark and AWS - Vice President

Citi

Pune District

On-site

INR 2,800,000 - 5,000,000

Full time

14 days+
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Job summary

Citi is seeking a Senior Databricks Engineer to advance the data processing platform on Databricks on AWS, helping migrate from legacy Cloudera Hadoop to cloud-native Databricks. The role focuses on Spark design, performance tuning, and modernizing pipelines while shaping scalable, production-ready solutions across teams.

The ideal candidate will lead complex implementation efforts, optimize Delta Lake usage, and collaborate with architects, DevOps, and platform teams to deliver robust data

Qualifications

  • 10+ years in data engineering or distributed systems.
  • Strong Spark, Databricks on AWS, and Delta Lake experience.
  • Proficient in SQL and AWS cloud services.
  • Experience modernizing legacy platforms to cloud-native architectures.

Responsibilities

  • Platform Engineering & Modernization: Refactor Spark pipelines for Databricks native architectures and eliminate Hadoop dependencies.
  • Databricks Native Development: Build and optimize using Delta Lake and Databricks Workflows for orchestration and auto scaling.
  • Design & Solution Engineering: Translate architecture into detailed designs and define data models and pipeline patterns.
  • Performance Optimization & Simplification: Improve Spark performance and simplify complex pipelines into efficient patterns.
  • Engineering Standards & Best Practices: Write clean, modular, testable code and contribute to shared frameworks.
  • Collaboration & Stakeholder Engagement: Work with senior architects, platform teams, and DevOps engineers.
  • Testing & Quality Assurance: Develop unit/integration tests and support production releases.

Skills

Spark (Java/PySpark)
Databricks on AWS
Delta Lake
SQL
AWS services
Data modeling

Education

Bachelor’s degree

Job description

We are looking for a highly skilled Senior Databricks Engineer to contribute to the engineering, modernization, and continuous evolution of data processing platform on Databricks on AWS. While supporting the transition from the legacy Cloudera Hadoop platform to Databricks on AWS, this role will continue to play a key part in enhancing performance, simplifying pipelines, and delivering new capabilities on the Databricks platform over the long term.

The ideal candidate is a strong hands‑on Spark engineer with solid design experience, capable of contributing to architectural decisions while leading complex implementation and optimization efforts.

Responsibilities
1. Platform Engineering & Modernization
  • Refactor and modernize existing Spark pipelines to Databricks native architectures
  • Eliminate legacy Hadoop dependencies and adopt cloud native AWS patterns
  • Enhance and extend existing processing logic using optimized Spark (JavaSpark / PySpark) on Databricks
2. Databricks Native Development
  • Build and optimize solutions using Databricks features, including Delta Lake, Databricks Workflows for orchestration and Auto scaling and job clusters
3. Design & Solution Engineering
  • Contribute to low and mid level architecture and design
  • Translate high level architecture into detailed technical designs
  • Define data models, pipeline patterns, and reusable components
  • Ensure solutions are scalable, maintainable, and production ready
4. Performance Optimization & Simplification
  • Analyze, improve Spark job performance and simplify complex or over engineered pipelines into standardized, efficient patterns
5. Engineering Standards & Best Practices
  • Follow and contribute to Databricks and Spark engineering standards
  • Write clean, modular, and testable code
  • Contribute to shared frameworks, reusable libraries, and quality standards
6. Collaboration & Stakeholder Engagement
  • Work closely with senior architects, platform teams, and DevOps engineers
  • Provide technical inputs, troubleshooting support, and implementation guidance
  • Participate in design discussions and technical decision making
7. Testing & Quality Assurance
  • Develop unit, integration, and data validation tests
  • Support production releases and post deployment validation
Qualifications
Core Technical Skills
  • 10+ years in data engineering or distributed systems
  • Strong expertise in Apache Spark (JavaSpark / PySpark), Databricks on AWS, and Delta Lake
  • Experience on SQL
  • Experience with AWS services and large‑scale distributed data processing
Modernization & Optimization Experience
  • Experience modernizing or refactoring legacy data platforms into cloud‑based architectures
  • Strong background in Spark performance tuning and large‑scale batch optimization
Design Capability
  • Ability to translate architecture into implementable designs
  • Understanding of data modeling and pipeline orchestration patterns
Behavioral Competencies
  • Strong problem‑solving mindset for complex distributed systems
  • Comfortable working in time‑bound, high‑impact environments
  • Proactive, accountable, and collaborative
  • Clear communication skills across global teams
Education
  • Bachelor’s degree/University degree or equivalent experience
Job Family Group

Technology

Job Family

Applications Development

Time Type

Full time

Most Relevant Skills

Please see the requirements listed above.

Other Relevant Skills

For complementary skills, please see above and/or contact the recruiter.

Citi is an equal opportunity employer, and qualified candidates will receive consideration without regard to their race, color, religion, sex, sexual orientation, gender identity, national origin, disability, status as a protected veteran, or any other characteristic protected by law.

If you are a person with a disability and need a reasonable accommodation to use our search tools and/or apply for a career opportunity review Accessibility at Citi.

View Citi’s EEO Policy Statement and the Know Your Rights poster.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Databricks Apache Spark - SQL Engineer- Vice President
Senior Databricks Apache Spark - SQL Engineer- Vice President

Citi • Pune District

On-site
INR 400,000 - 650,000
Senior Databricks Engineer, Apache Spark and AWS - Vice President
Senior Databricks Engineer, Apache Spark and AWS - Vice President

Citi • Maharashtra

On-site
INR 2,500,000 - 5,000,000
Senior Data Engineer - Apache Spark and SQL - Vice President
Senior Data Engineer - Apache Spark and SQL - Vice President

Citigroup Inc. • Pune District

On-site
INR 4,000,000 - 8,000,000
Senior Data Engineer - Apache Spark and SQL - Vice President
Senior Data Engineer - Apache Spark and SQL - Vice President

Citi • Pune District

On-site
INR 3,000,000 - 6,000,000
Senior Databricks Engineer, Apache Spark and AWS - Vice President
Senior Databricks Engineer, Apache Spark and AWS - Vice President

Citigroup Inc. • Pune District

On-site
INR 4,500,000 - 7,500,000
Senior Databricks Engineer, Apache Spark and AWS - Vice President
Senior Databricks Engineer, Apache Spark and AWS - Vice President

Citibank (Switzerland) AG • Pune District

On-site
Confidential
Senior Databricks Apache Spark -SQL Engineer - Assistant Vice President
Senior Databricks Apache Spark -SQL Engineer - Assistant Vice President

Citi • Pune District

On-site
INR 2,500,000 - 4,000,000
Senior Software Engineer - Core Java & Apache Spark
Senior Software Engineer - Core Java & Apache Spark

Citi • Chennai District

On-site
INR 1,800,000 - 2,500,000
Senior Data Engineer - Assistant Vice President
Senior Data Engineer - Assistant Vice President

Citibank (Switzerland) AG • Chennai District

On-site
Confidential
Senior Software Engineer - Core Java & Apache Spark
Senior Software Engineer - Core Java & Apache Spark

Citigroup Inc. • Chennai

On-site
INR 1,000,000 - 1,500,000