Lead Software Engineer - Data Engineer + Python + Spark/Pyspark+ AWS

JPMorgan Chase & Co.

Bengaluru

On-site

INR 4,000,000 - 7,000,000

Full time

13 days ago

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

JPMorgan Chase & Co. in Bengaluru invites a Lead Software Engineer - Data Engineering to architect and deliver reliable, scalable data platforms powering analytics and AI-assisted workflows.

You will own data contracts, lineage, and SLAs while ensuring governance, security, and resilience across Java and Python implementations. You will partner with product, UX, and platform teams to enable AI-enabled capabilities, develop robust ETL/ELT processes, and lead orchestration, testing, and production

Qualifications

  • Formal training or certification in software engineering and 5+ years of applied experience.
  • Hands-on production-grade data platforms and pipelines delivery, including leadership.
  • Strong Java and Python production proficiency, with performance-minded coding.
  • Advanced SQL and data modeling skills, with experience in orchestration (Airflow).
  • Experience with dbt, ETL/ELT, testing, and data quality practices.
  • Proven ability to design scalable batch/stream processing and search/indexing worklows.

Responsibilities

  • Design, build, and operate batch and streaming data pipelines with reliability and cost-efficiency.
  • Lead data modeling for domain datasets, including contracts, lineage, and SLAs.
  • Develop ETL/ELT workflows with validation controls and alerting.
  • Engineer data processing using Java and Python across platforms and services.
  • Build orchestration (Airflow) with scheduling, retries, and end-to-end ownership.
  • Deliver transformation pipelines using dbt with strong testing and release discipline.
  • Develop and maintain APIs exposing curated data products for downstream use.
  • Build and maintain search/indexing workflows (Elasticsearch) with quality standards.
  • Collaborate with security, risk, and controls to meet governance needs.
  • Advance AI/ML readiness with governance, guardrails, and traceability.

Skills

Java
Python
SQL
Data modeling
ETL/ELT development
CI/CD
Security & governance
Communication

Tools

Airflow
dbt
Elasticsearch
Kafka
Spark
Spring Batch
Kubernetes
AWS

Job description

  • Job Identification 210752829
  • Job Category Software Engineering
  • Business Unit Commercial & Investment Bank
  • Posting Date 08/26/2026, 08:32 PM
  • Locations Parcel 9, Embassy Tech Village, Outer Ring Road, Deverabeesanhalli Village, Varthur Hobli, Bengaluru, IN-KA, 560103, IN
  • Apply Before 09/05/2026, 04:00 AM
  • Job Schedule Full time
  • Job Shift Day
Job Description


We have an opportunity to impact your career and provide an adventure where you can push the limits of what's possible.


As a Lead Software Engineer - Data Engineering at JPMorgan Chase within Global Banking Technology, you will lead the design and delivery of reliable, scalable data platforms and pipelines that power business-critical use cases, including analytics, search/retrieval, and AI-assisted workflows. You will be accountable for building high-quality curated datasets with clear contracts, lineage expectations, and measurable SLAs/SLOs, while ensuring strong controls across security, privacy, resiliency, and auditability. You will remain hands-on and will set the technical bar for engineering rigor across Java and Python implementations, including batch/stream processing, microservices, and shared libraries. You will also partner with product, UX, and platform teams to enable AI/ML and agentic patterns where they materially improve business outcomes, without compromising governance or operational discipline.

Job Responsibilities

  • Design, build, and operate batch and streaming data pipelines that are reliable, observable, and cost-efficient, with clear runbooks and production support ownership.
  • Lead data modeling and curation for domain datasets, including schema evolution, data contracts, lineage expectations, and consumer-facing SLAs/SLOs.
    Implement robust ETL/ELT workflows with strong validation controls, including reconciliation, completeness checks, anomaly detection, and automated alerting.
  • Engineer high-throughput data processing solutions using a combination of Java and Python, selecting the right tool for performance, maintainability, and platform standards.
  • Build and operate orchestration capabilities (for example, Airflow or equivalent), including scheduling, backfills, retries, dependency management, and operational SLAs with end-to-end ownership across architecture, engineering standards, CI/CD, and operational stability in a regulated enterprise context.
  • Deliver transformation pipelines using modern transformation frameworks (for example, dbt or equivalent), with strong testing, repeatability, and release discipline.
  • Develop and maintain supporting services and APIs (REST and/or gRPC) that expose curated data products and enable downstream consumers, using clean architecture and well-defined contracts.
  • Build and maintain search and indexing pipelines (for example, Elasticsearch) that support discovery, retrieval, analytics, and RAG-style experiences, and establish engineering and data quality standards across code review, automated testing, performance tuning, observability (logs/metrics/traces), resiliency patterns, and incident response.
  • Partner with security, risk, and controls teams to ensure data solutions meet governance expectations, including access controls, secrets handling, least privilege, and auditability.
  • Enable secondary AI/ML capabilities by delivering the data foundations required for evaluation, guardrails, tool/function integration, and traceable AI-assisted workflows, including MCP-style integration patterns where applicable.
  • Drives team adoption of enterprise-authorized AI-assisted engineering practices within the work environment to improve code quality, delivery speed, and operational outcomes (e.g., AI-assisted code review/refactoring, test strategy acceleration, incident/root-cause analysis support), while establishing consistent validation standards (secure coding, peer review, automated testing) and promoting reuse of effective patterns across the team.
  • Applies knowledge of tools within the Software Development Life Cycle toolchain, including enterprise-authorized AI-assisted development and automation capabilities, to improve the value realized by automation.

Required qualifications, capabilities, and skills

  • Formal training or certification on software engineering concepts and 5+ years applied experience
  • Hands-on engineering experience delivering production-grade platforms and data systems, with demonstrated recent experience as a lead Data Engineer building and operating curated datasets and production pipelines end-to-end.
  • Strong hands-on proficiency in both Java and Python in production environments, including performance-minded development, design patterns, and maintainable codebases.
  • Advanced SQL skills, with strong capability in data modeling, schema design, and schema evolution. with proven experience with pipeline orchestration (for example, Airflow or equivalent), including operational controls (SLAs, alerting, retries, backfills).
  • Proven experience with transformation frameworks (for example, dbt or equivalent) and strong testing practices for transformations and data quality.
  • Experience processing large-scale datasets with a clear track record of optimizing for performance, scalability, reliability, and cost.
  • Experience designing and implementing large-scale batch processing jobs (for example, Spring Batch or equivalent enterprise batch frameworks).
  • Hands-on experience building and operating search/indexing workflows (for example, Elasticsearch) at scale with strong SDLC discipline: code reviews, unit/integration testing, CI/CD, release hygiene, and production support ownership.
  • Secure engineering fundamentals: authentication/authorization, secrets management, least privilege, secure coding, and policy enforcement patterns (including familiarity with OPA or similar policy-as-code approaches) with strong communication and cross-functional leadership across engineering, product, UX, platform, and control partners.
  • Demonstrated experience leading effective use of approved AI-assisted software development tools (e.g., for coding, code review, test acceleration, troubleshooting) with the ability to set team expectations for validating AI outputs for correctness, performance, and security.
  • Strong understanding of responsible AI use in engineering workflows, including data sensitivity considerations, secure handling of inputs/outputs, and adherence to resiliency and security expectations; experience coaching engineers on safe, compliant adoption within delivery practices

Preferred qualifications, capabilities, and skills

  • Experience with event streaming and asynchronous architectures (for example, Kafka) and event-driven processing patterns.
  • Experience with large-scale processing engines (for example, Spark or equivalent) and distributed compute cost governance.
  • Cloud-native delivery experience (for example, AWS), including containers and Kubernetes, with strong operational excellence practices.
  • Experience building LLM/GenAI-enabled applications, including RAG patterns, evaluation approaches, and safety controls, with a disciplined approach to governance and traceability.
  • Familiarity with agentic architectures, including orchestrators, tool/function integrations, workflow/state management, and MCP-style integration concepts.
  • Experience delivering in regulated environments with strong risk, control, and audit requirements.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Sr Lead Software Engineer - Java/Python, Data Engineering
Sr Lead Software Engineer - Java/Python, Data Engineering

JPMorgan Chase & Co. • Hyderabad

On-site
INR 3,000,000 - 5,400,000
Sr Lead Software Engineer - Java/Python, Data Engineering
Sr Lead Software Engineer - Java/Python, Data Engineering

JPMorganChase • Hyderabad

On-site
INR 4,000,000 - 7,000,000
Lead Software Engineer- DevOps, Kubernetes, AWS
Lead Software Engineer- DevOps, Kubernetes, AWS

JPMorgan Chase & Co. • Bengaluru

On-site
INR 1,400,000 - 2,800,000
Lead Software Engineer - Lead Data Architect
Lead Software Engineer - Lead Data Architect

JPMorganChase • Mumbai

On-site
INR 4,000,000 - 7,000,000
Lead Software Engineer - Data Platforms, REST APIs, SQL
Lead Software Engineer - Data Platforms, REST APIs, SQL

JPMorganChase • Hyderabad

On-site
INR 4,200,000 - 6,800,000
Lead Software Engineer
Lead Software Engineer

JPMorganChase • Bengaluru

On-site
INR 400,000 - 700,000
Lead Software Engineer - Python, Data, Cloud
Lead Software Engineer - Python, Data, Cloud

JPMorganChase • Bengaluru

On-site
INR 4,000,000 - 7,000,000
Sr Lead Software Engineer - Java Python Data Engineering
Sr Lead Software Engineer - Java Python Data Engineering

JPMorganChase • Hyderabad

On-site
INR 300,000 - 600,000
Senior Lead Software Engineer - Java/Python, React, AWS
Senior Lead Software Engineer - Java/Python, React, AWS

JPMorgan Chase & Co. • Bengaluru

On-site
INR 4,000,000 - 6,500,000
Lead Software Engineer - Java and Python Fullstack Developer
Lead Software Engineer - Java and Python Fullstack Developer

JP Morgan Services India Pvt Ltd • Bengaluru

On-site
INR 3,500,000 - 5,500,000