Data Engineer - Python, AI

Citi

Maharashtra

On-site

INR 2,400,000 - 4,000,000

Full time

4 days ago
Be an early applicant
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Job summary

Citi in India seeks a mid-level Python Developer blending data engineering and AI/NLP work to build scalable NLP pipelines and data platforms. You will develop PySpark data jobs, REST APIs with Flask, and MLflow-based model tracking, while collaborating across teams.

You should be proficient with PyTorch/TensorFlow optional, Redis caching, CI/CD via GitHub, and Linux shell scripting, with a strong grasp of design patterns and debugging.

Qualifications

  • 10–12 years of hands-on Python programming experience.
  • Strong Python, OOP and design patterns fundamentals.
  • Experience with NLP libraries such as Flair, BERT, Transformers.
  • Solid PySpark, Pandas, PyArrow and distributed data pipelines.
  • APIs using Flask; CI/CD and Git workflows experience.
  • Redis in-memory store experience.
  • Autosys JILs for job scheduling experience.
  • Linux command line and shell scripting.
  • Debugging, problem-solving and teamwork skills.
  • Cloud exposure; AWS boto3 experience is an asset.

Responsibilities

  • Develop and optimize ETL/data processing jobs using PySpark and Pandas.
  • Build NLP pipelines using Flair, BERT, and LLM-based models.
  • Create scalable ingestion and data transformation pipelines for AI use cases.
  • Develop Flask-based APIs for model inference and service integrations.
  • Use regular expressions for text cleaning and NLP preprocessing.
  • Integrate Redis caching and fast lookups.
  • Manage ML models with MLflow for tracking and versioning.
  • Support CI/CD workflows with GitHub and LightSpeed Enterprise.
  • Create and maintain Autosys JILs for scheduling.
  • Write unit tests with PyTest/unittest.

Skills

Python programming
OOP & design patterns
NLP libraries
PySpark
Pandas
Flask API
MLflow
CI/CD
Redis
Linux CLI
REST APIs

Tools

Docker
Kubernetes
FastAPI
Airflow
Prefect
PyTorch
TensorFlow
boto3

Job description

Role Summary

We are looking for a mid-level Python Developer with combined experience in Data Engineering and AI/NLP engineering. The candidate will build NLP pipelines using libraries such as Flair, BERT, and LLM frameworks, and will also work on large-scale data processing using PySpark, Pandas, and related data tools. The role includes developing APIs, integrating with platform services, and supporting CI/CD deployments using GitHub and LightSpeed Enterprise.

Key Responsibilities
  • Develop and optimize ETL/data processing jobs using PySpark, Pandas, PyArrow, and related libraries.
  • Build and maintain NLP pipelines using Flair, BERT, and LLM-based models.
  • Develop scalable ingestion and data transformation pipelines for AI and analytics use cases.
  • Build and maintain Flask-based APIs for model inference and service integrations.
  • Use regular expressions for text cleaning, parsing, and NLP preprocessing.
  • Integrate caching and fast lookups using Redis.
  • Manage and deploy ML models using MLflow for tracking and versioning.
  • Support CI/CD workflows using GitHub, LightSpeed Enterprise, and deployment pipelines.
  • Create and maintain Autosys JILs for job scheduling and automation.
  • Use basic Linux commands for troubleshooting, operations, and deployment tasks.
  • Monitor application and system health using ITRS Geneos.
  • Write unit tests and improve automation test coverage (PyTest/unittest).
  • Work with REST APIs, microservices, and basic shell scripting.
  • Work with cloud services (ECS), including boto3.
Required Skills
  • 10–12 years of hands‑on Python programming experience.
  • Strong fundamentals in Python, OOP, and design patterns.
  • Experience with NLP libraries such as Flair, BERT, HuggingFace Transformers, or similar.
  • Solid experience with PySpark, Pandas, PyArrow, and distributed data pipelines.
  • Experience building APIs using Flask (FastAPI is a plus).
  • Experience with MLflow for model tracking and deployment.
  • Good understanding of CI/CD practices and Git workflows.
  • Experience working with Redis or similar in‑memory stores.
  • Experience with Autosys JILs for job scheduling.
  • Comfortable with Linux command line and shell scripting.
  • Strong debugging, problem‑solving, and teamwork skills.
  • Exposure to cloud services; AWS boto3 experience is an asset.
Nice-to-Have
  • Experience with Polars or Dask for high-performance data processing.
  • Experience with PyTorch or TensorFlow for model training.
  • Experience with Docker, Kubernetes, or containerized deployments.
  • Experience with monitoring tools such as ITRS Geneos.
  • Experience with FastAPI, Airflow, or Prefect.

------------------------------------------------------

Job Family Group

Technology

------------------------------------------------------

Job Family

Applications Development

------------------------------------------------------

Time Type

Full time

------------------------------------------------------

Most Relevant Skills

Please see the requirements listed above.

------------------------------------------------------

Other Relevant Skills

For complementary skills, please see above and/or contact the recruiter.

------------------------------------------------------

Citi is an equal opportunity employer, and qualified candidates will receive consideration without regard to their race, color, religion, sex, sexual orientation, gender identity, national origin, disability, status as a protected veteran, or any other characteristic protected by law.

If you are a person with a disability and need a reasonable accommodation to use our search tools and/or apply for a career opportunity review Accessibility at Citi.

View Citi’s EEO Policy Statement and the Know Your Rights poster.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Data Engineer - Python, AI
Data Engineer - Python, AI

Citibank (Switzerland) AG • Pune District

Hybrid
Confidential
Senior Python Data and AI Engineer - Vice President
Senior Python Data and AI Engineer - Vice President

Citigroup Inc. • Pune District

On-site
INR 3,500,000 - 6,000,000
Data Engineer - Big Data Python
Data Engineer - Big Data Python

Citi • Chennai District

On-site
INR 1,200,000 - 2,100,000
Python Engineering AI Lead-Assistant Vice president
Python Engineering AI Lead-Assistant Vice president

Citi • Chennai District

On-site
INR 3,000,000 - 6,000,000
Big data/Python/Databricks Engineer Engineer
Big data/Python/Databricks Engineer Engineer

Citi • Chennai District

Hybrid
INR 1,500,000 - 2,500,000
Hybrid work model
Continuous learning resources
Structured career pathway
Python Engineering AI Lead-Assistant Vice president
Python Engineering AI Lead-Assistant Vice president

Citibank (Switzerland) AG • Chennai District

Hybrid
Confidential
Senior Developer - Python and Spark
Senior Developer - Python and Spark

Citigroup Inc. • Pune District

Hybrid
INR 2,800,000 - 4,200,000
Hybrid work arrangement
Exposure to AI/data platform projects
Growth in AI and platform architecture
Senior Python Application Developer
Senior Python Application Developer

Citigroup Inc. • Pune District

On-site
INR 1,500,000 - 2,500,000
Data Engineer Python Developer
Data Engineer Python Developer

Citi • Pune District

On-site
INR 1,200,000 - 1,800,000
Senior Developer - Python and Spark
Senior Developer - Python and Spark

Citi • Maharashtra

Hybrid
INR 1,200,000 - 1,800,000
Hybrid work schedule