Build secure, scalable data engineering solutions in a role focused on modernizing to a Databricks-on-AWS lakehouse. At JPMorganChase, you’ll design and deliver production-grade pipelines using Apache Spark (PySpark), optimize for performance and cost, and help enable trusted analytics for business partners. The team also emphasizes a culture of diversity, opportunity, inclusion, and respect, with support for engineering growth through communities of practice and emerging technologies.
What you’ll do
- Design, build, and maintain new data pipelines on Databricks using PySpark, delivering secure, high-quality production code and collaborating through reviewing and debugging work by others.
- Optimize PySpark jobs and Databricks clusters for performance, scalability, and cost efficiency, applying techniques such as partitioning, caching, and resource management.
- Create scalable data frameworks for end-to-end Databricks pipelines for workforce data analytics using medallion (bronze/silver/gold) lakehouse patterns.
- Implement data quality checks and validation using Delta Lake and Delta Live Tables expectations to support accurate, reliable data.
- Build monitoring and alerting to proactively address data ingestion issues, leveraging Databricks and AWS CloudWatch to improve performance and throughput.
- Find opportunities to eliminate or automate remediation for recurring issues using Databricks Workflows and AWS-native automation.
- Use AI and Agentic AI to accelerate pipeline development, including adopting AI-assisted engineering tools such as Claude and GitHub Copilot.
- Provision and deliver curated, reliable datasets for BI partners using Sigma, Tableau, and Alteryx to support reporting and analytics.
- Partner with stakeholders to understand requirements and design solutions, producing architecture and design artifacts for complex applications.
- Contribute to software engineering communities of practice that explore new and emerging technologies, reinforcing a strong culture of inclusion and respect.
What you bring
- 3+ years of applied experience in data engineering, with formal training or certification in software engineering concepts.
- Advanced, hands‑on expertise in Apache Spark (PySpark) for large‑scale distributed processing, plus strong experience building and operating production pipelines on Databricks with Delta Lake and lakehouse patterns.
- Strong knowledge of the AWS data ecosystem, including S3, EMR, Glue, Lambda, and Athena, along with AWS storage and compute services; experience with Parquet and Iceberg formats.
- Strong programming skills in Python (Java or Scala is a plus).
- Experience with automation and continuous delivery, using CI/CD pipelines and tools such as Git/Bitbucket, Jenkins, or Spinnaker.
- Hands‑on experience spanning system design, application development, testing, and operational stability, including agile methodologies, resiliency, and security.
- In‑depth knowledge of the financial services industry and its IT systems.
- Solid SQL and data modeling skills, including experience with relational databases such as Oracle (a plus).
- Practical experience with scheduling tools such as Airflow and Autosys.
Technologies you’ll work with
Databricks, Apache Spark, PySpark, AWS, Delta Lake, Delta Live Tables, Databricks Workflows, AWS CloudWatch, AI, Agentic AI, Claude, GitHub Copilot, Sigma, Tableau, Alteryx, Python, Java, Scala, Git, Bitbucket, Jenkins, Spinnaker, S3, EMR, Glue, Lambda, Athena, Parquet, Iceberg, SQL, Oracle, Airflow, Autosys.
Benefits
- Competitive total rewards package including base salary determined based on role, experience, skill set, and location.
- In eligible roles, commission‑based pay and/or discretionary incentive compensation paid as cash and/or forfeitable equity.
- Comprehensive health care coverage.
- On‑site health and wellness centers.
- Retirement savings plan.
- Backup childcare.
- Tuition reimbursement.
- Mental health support.
- Financial coaching.
- Additional details about total compensation and benefits provided during the hiring process.
- Equal opportunity employer with high value on diversity and inclusion.
Preferred qualifications
- Databricks certifications (for example: Databricks Certified Data Engineer Associate/Professional).
- Familiarity with Generative AI and Agentic AI frameworks, including experience using Claude and GitHub Copilot in an engineering workflow.
- Deeper expertise in the AWS cloud platform and its broader service catalog.
Team context
JPMorganChase professionals in Corporate Functions span a diverse set of areas from finance and risk to human resources and marketing. Corporate teams play an essential role in helping set businesses, clients, customers, and employees up for success.