Data Engineer III - Databricks, Pyspark, Python, AWS

Next Frontier Capital

Bengaluru

On-site

INR 2,400,000 - 3,400,000

Full time

6 days ago
Be an early applicant
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Job summary

JPMorgan Chase invites a Data Engineer III to join an agile team that designs and delivers trusted data pipelines and analytics solutions using Databricks, PySpark, Python, and AWS. You will develop, test, and maintain scalable data architectures across multiple business functions.

The role emphasizes data modelling, SQL optimization, and batch/streaming processing with a focus on reliability, performance, and security.

Qualifications

  • Formal training or certification on data engineering concepts and 3+ years applied experience.
  • Expert-level SQL skills with complex joins, window functions, CTEs, optimization, and analytics at scale.
  • Proficient Python in production environments with maintainable code.
  • Deep expertise in Apache Spark and PySpark, including performance tuning.
  • Strong Databricks experience for large-scale data processing.
  • Solid data warehousing concepts, data modelling, and scalable architecture.
  • Experience with batch and streaming data processing.
  • Familiarity with AWS S3 and common AWS data services.
  • Excellent debugging, troubleshooting, and version control (GitHub/Bitbucket).
  • Ability to validate AI-assisted outputs and ensure data handling compliance.

Responsibilities

  • Design, develop, and maintain big data pipelines (batch and streaming) using PySpark/Spark and Databricks.
  • Lead data modelling and solution design for data products, defining target-state architecture and data flows.
  • Build scalable ingestion and transformation workflows for high-volume datasets ensuring reliability and quality.
  • Develop and optimize SQL transformations and analytical datasets with query tuning.
  • Apply data warehousing concepts to build analytics-ready data layers; leverage AWS services like S3.
  • Perform advanced debugging across distributed Spark workloads (data skew, shuffle tuning, memory optimization).
  • Adopt modular design, clean coding, code reviews, CI-friendly development.
  • Use GitHub/Bitbucket for version control and releases.
  • Collaborate with cross-functional stakeholders to translate requirements into data solutions.
  • Review AI-assisted outputs for safety and accuracy before use.

Skills

SQL (expert level)
Python
PySpark
Distributed processing
Data modelling
Spark
Databricks
AWS
CI/CD
Team collaboration

Education

Data engineering certification

Tools

Databricks
PySpark
Spark
AWS S3
GitHub/Bitbucket
Version control

Job description

Be part of a dynamic team where your distinctive skills will contribute to a winning culture and team. As a Data Engineer III - Databricks, Pyspark, Python, AWS at JPMorgan Chase within the Commercial & Investment Bank, you'll serve as a seasoned member of an agile team to design and deliver trusted data collection, storage, access, and analytics solutions in a secure, stable, and scalable way. You are responsible for developing, testing, and maintaining critical data pipelines and architectures across multiple technical areas within various business functions in support of the firm’s business objectives.

Job responsibilities
  • Design, develop, and maintain big data pipelines (batch and streaming) using PySpark/Spark and Databricks.
  • Lead/own data modelling and solution design for data products, including defining target-state architecture, data flows, and transformation patterns.
  • Build scalable ingestion and transformation workflows for high-volume datasets, ensuring reliability, quality, and performance.
  • Develop and optimize complex SQL transformations, reconciliation queries, and analytical datasets; perform query tuning for large-scale workloads.
  • Apply strong data warehousing concepts (dimensional modeling, SCDs, partitioning strategies, etc.) to build well-structured, analytics-ready data layers and Leverage common AWS services, with strong emphasis on S3 and AWS data processing capabilities, to support scalable storage and processing.
  • Perform advanced debugging and troubleshooting across distributed Spark workloads (data skew, shuffle tuning, memory/compute optimization).
  • Implement engineering best practices: modular design, efficient coding, code reviews, and CI-friendly development approaches.
  • Use GitHub/Bitbucket and standard version control workflows to manage codebase, peer reviews, and releases.
  • Partner with cross-functional stakeholders to convert requirements into robust big data solutions.
  • Uses enterprise-authorized AI capabilities within the work environment to accelerate data pipeline/design analysis and documentation, validating outputs and handling data according to sensitivity and security requirements.
  • Applies reuse-first, AI-assisted practices to strengthen SDLC-quality routines for data pipelines (e.g., test generation and control validation), ensuring traceability/auditability and alignment to resiliency and security expectations.

Required qualifications, capabilities, and skills
  • Formal training or certification on data engineering concepts and 3+ years applied experience
  • Experience in data engineering / big data engineering, with strong hands-on delivery and Expert-level SQL skills (must be extremely strong): complex joins, window functions, CTEs, optimization, and analytical problem solving at scale.
  • Strong hands-on coding experience with Python in production environments; demonstrated ability to write efficient, maintainable code.
  • Deep expertise in Apache Spark (in depth) and PySpark, including performance tuning and distributed processing fundamentals.
  • Strong experience with Databricks for large-scale data processing and pipeline development.
  • Strong understanding of data warehousing concepts and best practices; proven capability in data modelling and solution design for scalable, maintainable data platforms/products.
  • Experience implementing both batch and streaming data processing solutions.
  • Familiarity with AWS S3 and common AWS services used in data platforms and processing
  • Excellent debugging, troubleshooting, problem-solving skills and experience with GitHub, Bitbucket, and version control best practices.
  • Demonstrated experience using enterprise-authorized AI capabilities within the work environment to support data engineering workflows with strong validation habits and awareness of data sensitivity.
  • Ability to review and validate AI-assisted outputs (e.g., query suggestions, test ideas, or model change summaries) before use, escalating when uncertain and following data handling requirements.

Preferred qualifications, capabilities, and skills
  • Good to have: infrastructure provisioning in AWS using Infrastructure as Code (IaC) (e.g., Terraform, AWS CloudFormation).

JPMorganChase, one of the oldest financial institutions, offers innovative financial solutions to millions of consumers, small businesses and many of the world’s most prominent corporate, institutional and government clients under the J.P. Morgan and Chase brands. Our history spans over 200 years and today we are a leader in investment banking, consumer and small business banking, commercial banking, financial transaction processing and asset management.

We recognize that our people are our strength and the diverse talents they bring to our global workforce are directly linked to our success. We are an equal opportunity employer and place a high value on diversity and inclusion at our company. We do not discriminate on the basis of any protected attribute, including race, religion, color, national origin, gender, sexual orientation, gender identity, gender expression, age, marital or veteran status, pregnancy or disability, or any other basis protected under applicable law. We also make reasonable accommodations for applicants’ and employees’ religious practices and beliefs, as well as mental health or physical disability needs. Visit our FAQs for more information about requesting an accommodation.

J.P. Morgan’s Commercial & Investment Bank is a global leader across banking, markets, securities services and payments. Corporations, governments and institutions throughout the world entrust us with their business in more than 100 countries. The Commercial & Investment Bank provides strategic advice, raises capital, manages risk and extends liquidity in markets around the world. Develop, test, and maintain critical data pipelines and architectures across multiple technical areas

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Data Engineer III - Databricks, Pyspark, Python, AWS
Data Engineer III - Databricks, Pyspark, Python, AWS

JPMorgan Chase & Co. • Bengaluru

On-site
INR 3,500,000 - 5,500,000
Data Engineer III - Databricks, Pyspark, Python, AWS
Data Engineer III - Databricks, Pyspark, Python, AWS

Fairygodboss • Bengaluru

On-site
INR 1,500,000 - 2,300,000
Data Engineer III - Databricks, Pyspark, Python, AWS
Data Engineer III - Databricks, Pyspark, Python, AWS

JPMorganChase • Bengaluru

On-site
INR 2,400,000 - 4,000,000
Global Banking Finance Data and Transformation Associate
Global Banking Finance Data and Transformation Associate

Next Frontier Capital • Bengaluru

On-site
INR 1,500,000 - 2,300,000
Director Of Software Engineering
Director Of Software Engineering

JPMorganChase • Mumbai

On-site
INR 6,000,000 - 9,000,000
Data Scientist Lead
Data Scientist Lead

Next Frontier Capital • Bengaluru

On-site
INR 3,500,000 - 5,500,000
Global Banking Finance Data and Transformation Associate
Global Banking Finance Data and Transformation Associate

JPMorganChase • Bengaluru

On-site
INR 1,800,000 - 2,800,000
Data Scientist Lead
Data Scientist Lead

Next Frontier Capital • Mumbai

On-site
INR 4,200,000 - 6,800,000
Analyst - Software Engineer I - AI ML Java Python REACT
Analyst - Software Engineer I - AI ML Java Python REACT

JPMorganChase • Mumbai

On-site
INR 1,400,000 - 2,000,000
Data Management Lead
Data Management Lead

JP Morgan Services India Pvt Ltd • Bengaluru

On-site
INR 1,800,000 - 2,400,000