Lead Software Engineer - Databricks, ML, AWS

慨正橡扯

Plano (TX)

On-site

USD 160,000 - 210,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

JPMorgan Chase in Plano, TX seeks a Lead Software Engineer, Machine Learning and Cloud to drive data pipelines and analytics across multiple teams. You will lead architecture, set up Databricks clusters, and advance AI-assisted development practices in a secure, scalable cloud environment.

Responsibilities include building high-throughput pipelines, Delta Lake governance, and CI/CD for data projects, with a focus on performance optimization and reliability.

Qualifications

  • Formal training or certification on software engineering concepts and 5+ years applied experience.
  • 8+ years of professional software/data engineering experience, including substantial production work with Spark on Databricks or EMR.
  • Demonstrated experience leading effective use of approved AI-assisted software development tools with the ability to set team expectations for validating AI outputs.
  • Strong understanding of responsible AI use in engineering workflows, including data sensitivity considerations and secure handling.
  • Strong proficiency in Python and/or Java for data processing, platform tooling, and automation.
  • Hands-on Databricks expertise (Delta Lake, Unity Catalog, Workflows, SQL Warehouses).

Responsibilities

  • Lead architecture and delivery of high-throughput, low-latency data pipelines using Databricks and Spark.
  • Establish lakehouse patterns with Delta Lake and ensure performance at scale.
  • Drive adoption of enterprise AI-assisted engineering practices to improve quality, delivery speed, and operations.
  • Own Databricks cluster strategy, including autoscaling and configuration.
  • Design secure data ingestion and transformation frameworks with Databricks services.
  • Enforce data quality, lineage, and governance; embed validations into pipelines.
  • Build reusable libraries, frameworks, and APIs in Python/Java; oversee testing.
  • Implement CI/CD for data projects and infrastructure deployments.

Skills

Python
Java
Databricks
Apache Spark
Airflow
SQL
Security
CI/CD
Leadership
AI-assisted tooling

Education

Formal software engineering training

Tools

Databricks
EMR
Delta Lake
Unity Catalog
Airflow
Terraform
Git-based workflows
Repos/notebooks

Job description

We have an exciting and rewarding opportunity for you to take your software engineering career to the next level.

As a Lead Software Engineer, Machine Learning and Cloud at JPMorgan Chase within the Corporate Technology- Consumer & Community Bank Finance group, you are an integral part of an agile team that works to enhance, build, and deliver trusted market-leading technology products in a secure, stable, and scalable way. As a core technical contributor and lead, you are responsible for conducting critical technology solutions across multiple technical areas within various business functions in support of the firm’s business objectives.

Job Responsibilities:

  • Lead architecture and delivery of high-throughput, low-latency data pipelines using Databricks and Apache Spark (Core, SQL, Structured Streaming).
  • Establish lakehouse patterns with Delta Lake (ACID transactions, schema evolution, time travel, Z-ordering, compaction) and ensure performance at scale.
  • Drives team adoption of enterprise-authorized AI-assisted engineering practices within the work environment to improve code quality, delivery speed, and operational outcomes (e.g., AI-assisted code review/refactoring, test strategy acceleration, incident/root-cause analysis support), while establishing consistent validation standards (secure coding, peer review, automated testing) and promoting reuse of effective patterns across the team.
  • Applies knowledge of tools within the Software Development Life Cycle toolchain, including enterprise-authorized AI-assisted development and automation capabilities, to improve the value realized by automation.
  • Own Databricks cluster strategy and setup: runtime selection, autoscaling, driver/executor sizing, Spark configs, unit scripts, cluster policies, pools, and instance profiles.
  • Orchestrate jobs with Databricks Workflows; integrate with AWS eventing and orchestration as needed.
  • Design secure data ingestion and transformation frameworks leveraging Databricks services: Design delta or unmanaged tables, Create tasks for data, ingestion process, Create DAGs using Airflow to orchestrate creation of trusted and refined data.
  • Enforce data quality, lineage, and governance using Unity Catalog and/or Glue Catalog; embed expectations and validation into pipelines.
  • Drive Spark performance engineering: partitioning strategies, file sizing, AQE, broadcast joins, shuffle tuning, caching, spill/memory control, and job right-sizing to optimize cost.
  • Build reusable libraries, frameworks, and APIs in Python and/or Java; oversee unit, integration, and data validation testing.
  • Implement CI/CD for data projects (Git-based workflows), Terraform Infrastructure deployments environment promotion, and automated deployments; champion engineering standards and code reviews.

Required qualifications, capabilities, and skills:

  • Formal training or certification on software engineering concepts and 5+ years applied experience.
  • 8+ years of professional software/data engineering experience, including substantial production work with Spark on Databricks or EMR.
  • Demonstrated experience leading effective use of approved AI-assisted software development tools (e.g., for coding, code review, test acceleration, troubleshooting) with the ability to set team expectations for validating AI outputs for correctness, performance, and security.
  • Strong understanding of responsible AI use in engineering workflows, including data sensitivity considerations, secure handling of inputs/outputs, and adherence to resiliency and security expectations; experience coaching engineers on safe, compliant adoption within delivery practices
  • Strong proficiency in Python and/or Java for data processing, platform tooling, and automation.
  • Hands-on Databricks expertise (Delta Lake, Unity Catalog, Workflows, Repos/notebooks, SQL Warehouses).
  • Proven track record architecting and operating ETL/ELT pipelines (batch and streaming), with schema design/evolution, SLAs, and reliability engineering.
  • Deep skills in Spark performance tuning and Databricks cluster setup/optimization.
  • Strong SQL and analytics data modeling (dimensional/star schema; lakehouse best practices).
  • CI/CD and automation tooling for data (Git workflows, artifact management) and testing frameworks (pytest, JUnit).
  • Security-first mindset: roles/instance profiles, secret management, encryption-at-rest/in-transit, and network controls.

Preferred qualifications, capabilities, and skills:

  • Experience with Delta Live Tables and advanced governance (catalogs, grants, auditing) in Databricks.
  • AWS networking knowledge (VPC, subnets, routing, security groups) and data egress controls.
  • Experience with Terraform for Infra deployments
  • Cost optimization experience: autoscaling strategies, spot vs on-demand, auto-termination, storage layouts and compaction.
  • Observability for data systems (freshness/completeness metrics, lineage, SLAs, alerting).
  • Drive databricks performance tuning through liquid clustering or partitioning keys, familiarity with Airflow, Genie, Streamlit and React
  • Demonstrated leadership in code quality, reviews, testing strategy, CI/CD, and technical mentorship; excellent communication with stakeholders.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Lead Software Engineer - Databricks/Spark/AWS
Lead Software Engineer - Databricks/Spark/AWS

JPMorgan Chase & Co. • Kentucky

On-site
USD 130,000 - 160,000
Senior Manager of Software Engineering - Databricks, AWS
Senior Manager of Software Engineering - Databricks, AWS

JPMorgan Chase & Co. • Plano (TX)

On-site
USD 130,000 - 150,000
Lead Software Engineer-Big Data Python / Java , Databricks
Lead Software Engineer-Big Data Python / Java , Databricks

JPMorgan Chase & Co. • Houston (TX)

On-site
USD 140,000 - 170,000
Lead Software Engineer- Big Data Python /Java , Databricks
Lead Software Engineer- Big Data Python /Java , Databricks

JPMorgan Chase & Co. • Houston (TX)

On-site
USD 140,000 - 210,000
Lead Software Engineer - Data Engineer
Lead Software Engineer - Data Engineer

JPMorgan Chase & Co. • Jersey City (NJ)

On-site
USD 120,000 - 150,000
Lead Software Engineer - Platform Engineering Databricks
Lead Software Engineer - Platform Engineering Databricks

JPMorgan Chase & Co. • Jersey City (NJ)

On-site
USD 120,000 - 150,000
Senior Manager of Software Engineering - Databricks, AWS
Senior Manager of Software Engineering - Databricks, AWS

JPMorgan Chase • Plano (TX)

On-site
USD 120,000 - 160,000
Comprehensive health care coverage
Retirement savings plan
Tuition reimbursement
+1
Software Engineer III - Databricks
Software Engineer III - Databricks

JPMorgan Chase & Co. • Wilmington (DE)

On-site
USD 135,000 - 195,000
Lead Software Engineer - Python, Databricks and AWS
Lead Software Engineer - Python, Databricks and AWS

慨正橡扯 • Jersey City (NJ)

On-site
USD 180,000 - 240,000
Competitive compensation
Health benefits
Career growth opportunities
Senior Manager of Software Engineering - Databricks, AWS
Senior Manager of Software Engineering - Databricks, AWS

慨正橡扯 • Plano (TX)

On-site
USD 120,000 - 160,000
Comprehensive health care coverage
Retirement savings plan
Tuition reimbursement
+2