Databricks Data Platform Architect & Lead Engineer

Socket.dev

Plano (TX)

On-site

USD 150,000 - 210,000

Full time

6 days ago
Be an early applicant
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Job summary

JPMorgan Chase & Co. in Plano, TX, seeks a senior data engineering lead to architect and deliver high-throughput data pipelines on Databricks with Spark. You will optimize Delta Lake, Unity Catalog, and orchestrations with Airflow and Terraform, while enforcing data quality, security, and AI-assisted development standards.

You’ll build reusable data tooling in Python/Java, drive CI/CD, and mentor engineers to improve reliability and performance of large-scale analytics platforms.

Qualifications

  • Formal training or certification on software engineering concepts and 5+ years applied experience.
  • Advanced experience in software engineering and data engineering, including significant production delivery with Apache Spark on Databricks and/or AWS EMR.
  • Advanced hands-on Databricks expertise across Delta Lake, Unity Catalog, Workflows, Repos/notebooks, and SQL Warehouses, including cluster configuration and optimization.
  • Proven ability to architect, build, and operate reliable ETL/ELT data pipelines (batch and streaming), including schema design/evolution, SLAs, and reliability engineering practices.
  • Deep Spark performance tuning skills, with experience diagnosing bottlenecks and optimizing jobs for scalability, cost, and runtime efficiency.
  • Strong programming proficiency in Python and/or Java for data processing, platform tooling, and automation.
  • Strong SQL and analytics data modeling expertise, including dimensional/star schema design and Lakehouse best practices.
  • Demonstrated experience leading effective use of approved AI-assisted software development tools (coding, code review, test acceleration, troubleshooting), including setting team expectations and validation standards for correctness, performance, and security of AI outputs.
  • Strong responsible‑AI and security‑first engineering mindset, including data sensitivity awareness, secure handling of inputs/outputs, roles/instance profiles, secrets management, encryption at rest/in transit, network controls, and adherence to resiliency and security expectations; experience coaching teams on safe, compliant adoption within delivery practices.
  • Demonstrated experience leading effective use of approved AI-assisted software development tools (e.g., for coding, code review, test acceleration, troubleshooting) with the ability to set team expectations for validating AI outputs for correctness, performance, and security.
  • Strong understanding of responsible AI use in engineering workflows, including data sensitivity considerations, secure handling of inputs/outputs, and adherence to resiliency and security expectations; experience coaching engineers on safe, compliant adoption within delivery practices.

Responsibilities

  • Lead the architecture and delivery of high-throughput, low-latency data pipelines on Databricks using Spark, driving performance and reliability.
  • Establish and evolve Lakehouse patterns with Delta Lake to ensure scalable platforms.
  • Own Databricks cluster strategy and configuration, including autoscaling and tuning.
  • Orchestrate pipelines with Databricks Workflows and integrate with AWS services as needed.
  • Design secure ingestion and transformation frameworks using Delta or unmanaged tables and Airflow DAGs.
  • Enforce data quality, lineage, and governance via Unity Catalog and AWS Glue Catalog.
  • Drive Spark/Databricks performance engineering and cost optimization through partitioning, AQE, caching, and right-sizing.
  • Build reusable libraries/frameworks/APIs in Python and/or Java with strong test coverage.
  • Implement CI/CD for data projects using Git, Terraform, and automated releases; promote secure coding and peer review.

Skills

Spark on Databricks
Delta Lake
Unity Catalog
SQL
Python
Java
CI/CD
AI-assisted tooling
Security
Data modeling

Tools

Databricks
Delta Lake
Unity Catalog
Airflow
Terraform
Git
Spark
Python
Java

Job description

JPMorgan Chase & Co. in Plano, TX, seeks a senior data engineering lead to architect and deliver high-throughput data pipelines on Databricks with Spark. You will optimize Delta Lake, Unity Catalog, and orchestrations with Airflow and Terraform, while enforcing data quality, security, and AI-assisted development standards.

You’ll build reusable data tooling in Python/Java, drive CI/CD, and mentor engineers to improve reliability and performance of large-scale analytics platforms.

Get your free, confidential resume review.
or drag and drop your file here.