Senior Data Engineer - Vice President - Python Development

Citibank (Switzerland) AG

Pune District

Hybrid

INR 3,500,000 - 5,200,000

Full time

9 days ago
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Job summary

Citibank (Switzerland) AG in Pune is seeking a Senior Data Engineer, Vice President, to design and implement data foundations for Generative AI products. You will build production-grade Python pipelines, model data for low-latency AI access, and shape data architectures across Markets.

The role requires deep Python expertise, strong PostgreSQL and SQL skills, and experience with modern data tooling (Parquet, Spark, Airflow) in a high-performance environment. A hybrid work model is offered.

Qualifications

  • Extensive hands-on experience in data engineering, software engineering, or related discipline.
  • Deep practical expertise in Python and production-grade software.
  • Proven track record delivering large-scale data platforms or data-intensive apps.
  • Strong SQL skills and PostgreSQL experience.
  • Experience with data modelling, indexing, partitioning, and performance tuning.
  • Familiarity with columnar formats (Parquet) and storage strategies.
  • Experience with Spark, Dask, Polars or equivalent for large datasets.
  • Knowledge of ETL/ELT architecture, incremental processing, idempotency, and schema evolution.
  • Experience with data quality controls, lineage, monitoring, and alerting.
  • Strong software engineering fundamentals and CI/CD practices.
  • Experience deploying services on Kubernetes/OpenShift.
  • Understanding of data governance, security, entitlements, and access control.
  • Ability to operate in ambiguity and influence technical direction.
  • Strong communication and collaboration skills.

Responsibilities

  • Design and build data architecture for Generative AI products across Markets.
  • Develop scalable Python pipelines to ingest, transform, and enrich billions of records.
  • Create high-performance data-serving layers with low latency access.
  • Define storage strategies around PostgreSQL and Parquet.
  • Optimize ingestion, transformation, storage, indexing, and query performance.
  • Implement data quality controls, lineage, and schema evolution.
  • Build observability with metrics, logs, and alerts for data pipelines.
  • Collaborate with AI, software, and architecture teams; contribute to standards.
  • Develop tests and CI/CD pipelines for reliable delivery.
  • Work with security, entitlements, and governance requirements.

Skills

Python
SQL
PostgreSQL
Pandas
PyArrow
SQLAlchemy
NumPy
Apache Parquet
Apache Spark
Dask
Polars
Airflow
CI/CD
Kubernetes

Education

Bachelor's or Master's degree in Computer Science, Engineering, or related quantitative discipline

Tools

GitHub Copilot

Job description

## Senior Data Engineer - Vice President - Python DevelopmentApply: Hybrid: Pune Maharashtra India: Full time: Posted Today: 26975837**Technology**Join a small, high-impact engineering team in Citi Markets Technology building the data foundation for greenfield Generative AI products across asset classes. We are looking for an experienced Data Engineer who combines deep Python expertise with strong engineering judgement and a practical understanding of how to build reliable, high-performance data platforms.The role goes beyond constructing ETL pipelines. You will help define how billions of records from diverse Markets data sources are collected, validated, transformed, governed, and made available to production AI applications with consistently low retrieval latency.If you want to work on ambitious data engineering problems, shape a platform from the ground up, and help define how data is engineered for production AI in Markets, this is the role for you.**The Team**We are a fast-moving team specialising in Generative AI within Markets Technology. We build greenfield products that span multiple asset classes and solve real business problems using modern AI and data engineering approaches.Our applications depend on a robust and well-designed data foundation. That means creating pipelines and serving layers that can process billions of records while preserving data quality, provenance, security, and operational reliability.The team is still small enough for every engineer to have genuine influence. Data engineering is a core part of the product architecture, not a downstream support function.**The Role**As a Vice President in the team, you will be a hands-on senior engineer responsible for designing and building the data foundation for greenfield Generative AI products across Markets, including our conversational AI platform.You will develop production-grade data pipelines and services, primarily using Python, to ingest, validate, transform, enrich, and serve data from a wide range of internal and external sources. You will help create a curated, high-performance data-serving layer that enables low-latency retrieval by our AI and application services.You will contribute directly to code while also shaping architecture, engineering standards, data models, quality controls, and operational practices. We are looking for someone who can make pragmatic technology choices, challenge assumptions, and take long-term ownership of the platform.**What You’ll Do*** Design, build, and evolve the data architecture supporting Generative AI products across Markets* Develop scalable Python pipelines that ingest, validate, transform, enrich, and integrate billions of records from diverse sources* Build curated, queryable datasets and high-performance serving layers for low-latency application access* Design and optimise the application’s data storage strategies, initially centred on PostgreSQL and Parquet* Select suitable processing approaches for each workload, from efficient in-process and columnar processing to distributed frameworks where required* Optimise ingestion, transformation, storage, indexing, and query performance* Establish robust controls for data quality, reconciliation, lineage, schema evolution, idempotency, and recovery* Build production observability into data pipelines, including metrics, logging, alerting, and operational diagnostics* Design solutions that respect data classification, entitlements, access controls, and security requirements* Contribute directly to code, architecture reviews, technical standards, and the wider engineering direction of the team* Build automated tests and CI/CD pipelines that enable reliable and repeatable delivery* Collaborate closely with AI engineers, software engineers, architects, product partners, and Markets stakeholders* Use AI-assisted engineering tools, including Devin and GitHub Copilot, to improve development quality and productivity**What We’re Looking For*** Extensive hands-on experience in data engineering, software engineering, or a closely related discipline* Deep practical expertise in Python and the ability to build maintainable, production-grade software* A proven track record of designing and delivering large-scale data platforms or data-intensive applications* Strong SQL skills and extensive experience with relational databases, particularly PostgreSQL* Experience with data modelling, indexing, partitioning, query optimisation, and database performance tuning* Strong knowledge of the Python data ecosystem, including libraries such as pandas, PyArrow, SQLAlchemy, and NumPy* Practical experience working with columnar formats such as Apache Parquet and selecting efficient storage and serialisation strategies* Experience processing large-scale datasets using technologies such as Apache Spark, Dask, Polars, or equivalent frameworks* Strong understanding of ETL and ELT architecture, including incremental processing, idempotency, failure recovery, and schema evolution* Experience implementing automated data quality controls, reconciliation, lineage, monitoring, and operational alerting* Strong software engineering fundamentals, including design, testing, maintainability, code review, and CI/CD* Experience deploying and operating services on container platforms such as Kubernetes or OpenShift* Understanding of data governance and security practices, including encryption, masking, classification, entitlements, and fine-grained access control* The ability to operate in ambiguity, take ownership, and influence the technical direction of a product* Strong communication skills and a collaborative approach to engineering**What Makes This Role Different**This is not a role focused on maintaining legacy ETL jobs or moving data between systems without understanding how it will be used.You will help build the data foundation of a new generation of AI products in Markets. The engineering challenges include integrating complex datasets, processing billions of records, delivering consistently low retrieval latency, and meeting the quality, security, and reliability standards expected of production financial systems.The role offers the opportunity to influence the architecture from an early stage, work closely with AI and application engineers, and take genuine ownership of a platform that is central to the product.**Preferred Experience*** Experience with workflow orchestration technologies such as Apache Airflow, Dagster, or Prefect* Familiarity with Generative AI applications, retrieval architectures, LLMs, agentic systems, or structured evaluation* Experience building data foundations for search, retrieval, analytics, machine learning, or AI applications* Knowledge of financial instruments, trading concepts, and data structures within FX, Equities, or other capital markets domains* Experience working with temporal, reference, market, or transactional data* A bachelor’s or master’s degree in Computer Science, Engineering, or another relevant quantitative discipline, or equivalent professional experience
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior Data Engineer - Vice President - Python Development
Senior Data Engineer - Vice President - Python Development

Citi • India

On-site
INR 3,000,000 - 9,000,000
Python Microservices and Distributed Systems Lead - Vice President
Python Microservices and Distributed Systems Lead - Vice President

Citibank (Switzerland) AG • Pune District

On-site
Confidential
Senior Gen AI Engineer - Vice President
Senior Gen AI Engineer - Vice President

Citibank (Switzerland) AG • Pune District

On-site
Confidential
Python Engineering AI Lead - Assistant Vice President
Python Engineering AI Lead - Assistant Vice President

Citibank (Switzerland) AG • Pune District

Hybrid
INR 4,000,000 - 8,000,000
Senior Developer - Python and Spark
Senior Developer - Python and Spark

Citibank (Switzerland) AG • Pune District

On-site
Confidential
Hybrid work model (3 days in office, 2
Global engineering network
Wellbeing and work-life balance
Lead Python & SQL Application Developer - Vice President
Lead Python & SQL Application Developer - Vice President

Citibank (Switzerland) AG • Chennai District

On-site
Confidential
Python Engineering AI Lead-Assistant Vice president
Python Engineering AI Lead-Assistant Vice president

Citibank (Switzerland) AG • Chennai District

On-site
Confidential
Senior Python Engineer
Senior Python Engineer

Citibank (Switzerland) AG • Pune District

Hybrid
INR 2,800,000 - 4,200,000
AI Lead Data Engineer
AI Lead Data Engineer

Generac International • Pune District

On-site
INR 3,500,000 - 6,200,000
Python Microservices and Distributed Systems Lead - Vice President
Python Microservices and Distributed Systems Lead - Vice President

Citigroup Inc. • Pune District

On-site
INR 4,000,000 - 8,000,000