Data Engineer - Assistant Vice President

Citi

Maharashtra

On-site

INR 1,200,000 - 1,800,000

Full time

3 days ago
Be an early applicant
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Job summary

Citi is seeking a hands-on Data Engineer to design and deploy scalable data pipelines using Python, PySpark and Databricks. You will optimize performance across Spark clusters, implement Delta Lake governance, and collaborate with data science and BI teams.

Role requires cloud experience (AWS or GCP), strong SQL, and a track record of building robust data architectures within lakehouse environments.

Qualifications

  • Proficiency in Python and SQL with production-grade practices.
  • Deep hands-on experience with Apache Spark (PySpark) for large datasets.
  • 3+ years Databricks development including Delta Lake, DLT, Unity Catalog, and Workflows.
  • Experience deploying Databricks in AWS or GCP and cloud-native components (S3, IAM, EMR/BigQuery equivalents).
  • Strong data warehousing concepts including dimensional modeling and Medallion architecture.

Responsibilities

  • Act as SME on pipeline architecture, distributed computing, and cloud-native integrations.
  • Champion code quality, automated testing, CI/CD and data governance.
  • Design and deploy scalable batch/real-time data pipelines (ETL/ELT).
  • Manage Databricks workspaces and secure cloud integrations (IAM, secrets, storage).
  • Diagnose performance bottlenecks in Spark, SQL, and Databricks jobs; optimize storage.

Skills

Python Proficiency
PySpark & Spark SQL
Databricks Experience
Advanced SQL
Cloud (AWS/GCP)
Delta Lake/Unity Catalog
CI/CD Pipelines

Tools

Databricks
AWS
GCP
Delta Lake
Unity Catalog
GitHub Actions

Job description

Key Responsibilities
  • Act as a Subject Matter Expert (SME): Provide technical guidance on pipeline architecture, distributed computing, and cloud-native integrations.
  • Drive Best Practices: Champion code quality, comprehensive automated testing, CI/CD automation, and rigorous data governance standards.
  • Data Pipeline Architecture & Development: Design, develop, and deploy scalable batch and real-time end-to-end data pipelines (ETL/ELT) using Python, PySpark, Spark SQL, and Databricks(Must have).
  • Cloud Infrastructure Integration: Deploy and maintain Databricks workspaces on cloud environments (AWS or GCP). Manage secure integrations with cloud storage (S3/GCS), access controls (IAM), secrets management (Vault/KMS), and serverless query engines.
  • Performance Optimization & Tuning: Diagnose and resolve performance bottlenecks in Spark clusters, SQL queries, and Databricks jobs. Optimize storage layouts using Delta Lake properties (e.g., Z-Ordering, partitioning, and vacuuming).
  • Data Quality & Governance: Implement automated data validation frameworks, data quality monitoring, and metadata management solutions utilizing Databricks Unity Catalog to ensure strict compliance with internal data governance policies and external financial regulations (such as BCBS 239).
  • Technical Leadership & Mentorship: Act as a technical lead within an Agile/Scrum environment. Lead peer code reviews, enforce coding standards, and mentor junior and mid-level data engineers (C10/C11).
  • DevOps & CI/CD: Establish and maintain automated CI/CD pipelines (using Jenkins, GitLab, or GitHub Actions) for packaging and deploying data engineering artifacts (dbt, Spark jobs, Databricks workflows).
  • Collaboration: Partner with Data Science teams to operationalize machine learning models, and work with business intelligence developers to build efficient semantic layers for reporting.
Technical Qualifications (Must-Haves)
  • Programming Languages: Strong, production-grade proficiency in Python (including standard libraries, pandas, and testing frameworks like pytest) and advanced SQL (including window functions, CTEs, and query optimization).
  • Distributed Computing: Deep hands‑on experience with Apache Spark (PySpark) for processing multi-terabyte datasets in a distributed cluster environment.
  • Unified Lakehouse Platforms: Minimum 3 years of hands‑on experience developing within Databricks. Expert knowledge of Delta Lake ACID transactions, Delta Live Tables (DLT), Unity Catalog, and Databricks Workflows is required.
  • Cloud Platforms: Extensive experience deploying Databricks within either Amazon Web Services (AWS) or Google Cloud Platform (GCP). Proficiency in cloud-native components (AWS S3, EC2, IAM, EMR, Athena, Redshift OR GCP GCS, Compute Engine, IAM, Dataproc, BigQuery) is a strict requirement.
  • Data Modeling: Solid understanding of data warehousing concepts, including dimensional modeling (Star and Snowflake schemas), slow‑changing dimensions (SCDs), and Medallion (Bronze/Silver/Gold) architecture design.
Preferred Qualifications & Certifications
  • Databricks Certifications: Databricks Certified Data Engineer Professional or Databricks Certified Associate Developer for Apache Spark.
  • Cloud Certifications: AWS Certified Solutions Architect / AWS Certified Data Engineer, or Google Cloud Professional Data Engineer.
Professional Competencies & Soft Skills
  • Problem‑Solving: Exceptional analytical and troubleshooting skills to resolve complex performance and data consistency issues in distributed systems.
  • Communication: Excellent verbal and written communication skills. Ability to articulate complex technical architectures to non‑technical business stakeholders.
  • Collaborative Mindset: Proactive team player who thrives in a diverse, global, cross‑functional engineering environment.
  • Adaptability: Ability to prioritize work, pivot quickly in response to changing business requirements, and master new technologies as they emerge.

Thanks !

Job Family Group

Technology

Job Family

Applications Development

Time Type

Full time

Most Relevant Skills

Please see the requirements listed above.

Other Relevant Skills

For complementary skills, please see above and/or contact the recruiter.

Citi is an equal opportunity employer, and qualified candidates will receive consideration without regard to their race, color, religion, sex, sexual orientation, gender identity, national origin, disability, status as a protected veteran, or any other characteristic protected by law.

If you are a person with a disability and need a reasonable accommodation to use our search tools and/or apply for a career opportunity review Accessibility at Citi.

View Citi’s EEO Policy Statement and the Know Your Rights poster.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Data Engineer - Apache Spark and SQL - Vice President
Senior Data Engineer - Apache Spark and SQL - Vice President

Citi • Maharashtra

On-site
INR 3,500,000 - 7,000,000
Senior Databricks Engineer, Apache Spark and AWS - Vice President
Senior Databricks Engineer, Apache Spark and AWS - Vice President

Citigroup Inc. • Pune District

On-site
INR 4,500,000 - 7,500,000
Big data/Python/Databricks Engineer Engineer
Big data/Python/Databricks Engineer Engineer

Citigroup Inc. • Chennai District

Hybrid
INR 600,000 - 1,000,000
Hybrid working model
Learning resources
Career progression opportunities
Senior Data Engineer - Apache Spark and SQL - Vice President
Senior Data Engineer - Apache Spark and SQL - Vice President

Citi • Pune District

On-site
INR 3,000,000 - 6,000,000
Senior Databricks Engineer, Apache Spark and AWS - Vice President
Senior Databricks Engineer, Apache Spark and AWS - Vice President

Citibank (Switzerland) AG • Pune District

On-site
Confidential
Senior Data Engineer - Apache Spark and SQL - Vice President
Senior Data Engineer - Apache Spark and SQL - Vice President

Citigroup Inc. • Pune District

On-site
INR 4,000,000 - 8,000,000
Senior Databricks Engineer, Apache Spark and AWS - Vice President
Senior Databricks Engineer, Apache Spark and AWS - Vice President

Citi • Maharashtra

On-site
INR 2,500,000 - 5,000,000
Senior Data Engineer - Assistant Vice President
Senior Data Engineer - Assistant Vice President

Citigroup Inc. • Chennai District

On-site
INR 1,800,000 - 2,400,000
Big data/Python/Databricks Engineer Engineer
Big data/Python/Databricks Engineer Engineer

Citi • Chennai District

Hybrid
INR 1,500,000 - 2,500,000
Hybrid work model
Continuous learning resources
Structured career pathway
Senior Data Engineer (APAC Region)
Senior Data Engineer (APAC Region)

Anrgi Tech Private Limited • Pune District

On-site
INR 2,000,000 - 2,800,000