Senior Data Engineer

Prosum

Glendale (CA)

On-site

USD 150,000 - 210,000

Full time

2 hours ago
Be an early applicant
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Job summary

Prosum in Glendale, CA is seeking a Senior Data Engineer to own the Core Data platform on Databricks, delivering batch and streaming Spark pipelines, platform governance, and solution architecture across AWS, Kubernetes and Airflow. You will collaborate with technologists to turn data into actionable insights that drive growth.

You will design and implement data pipelines using PySpark/Scala, work with stakeholders, maintain Unity Catalog governance, and ensure pipeline reliability with CI/CD

Qualifications

  • 5+ years of data engineering experience developing and operating large-scale data pipelines.
  • Deep hands-on experience with Databricks and Apache Spark (batch and streaming).
  • Proficiency with Databricks platform tooling (API, SDK, CLI) for automation and governance.
  • Proficient in SQL with advanced performance tuning.
  • Hands-on production experience with Airflow (MWAA) for orchestrating data pipelines.

Responsibilities

  • Design, write, test, and deploy data pipelines using PySpark, Scala, SQL, Python.
  • Meet with stakeholders to gather requirements and translate them into scalable data platform solutions.
  • Understand Databricks platform and tooling to diagnose errors, audit activity, and automate updates.

Skills

Databricks
Advanced SQL
Data pipelines

Education

Bachelor's degree

Tools

Airflow (MWAA)
Docker
Kubernetes
Snowflake

Job description

As a Senior Data Engineer, you will be pivotal in transforming data into actionable insights. Collaborate with our dynamic team of technologists to develop cutting‑edge data solutions that drive innovation and fuel business growth. You will own and operate the Core Data platform on Databricks, delivering reliable batch and streaming Spark pipelines, platform governance, and solution architecture across AWS, Kubernetes, and Airflow. Your expertise will be essential in explaining Spark architecture, recommending best‑fit Databricks solutions to stakeholders, and optimizing our data‑driven decision‑making processes. If you're passionate about leveraging data to make a tangible impact, we welcome you to join us in shaping the future of our organization.

Key Responsibilities:
  • Design, write, test, and deploy data pipelines using PySpark, Scala, SQL, Python
  • Meet with stakeholders to gather requirements and translate them into scalable data platform solutions
  • Understanding of Databricks platform and developer tooling to diagnose errors, audit platform activity, and automate updates across pipelines, objects, and integrations
  • Ability to explain Spark architecture and pipeline behavior to stakeholders to diagnose root causes and recommend solutions
  • Provide solution architecture across AWS, Databricks, Kubernetes, and Airflow (MWAA), including cross‑platform integrations
  • Manage Databricks platform governance, including Unity Catalog, ACLs, lineage, and data discovery and privacy tooling
  • Build and maintain Kubernetes containers and containerized utilities supporting deployed data platform services
  • Apply networking knowledge to troubleshoot connectivity and integration errors across platform components
  • Perform platform administration: provision and remove access, assess resource utilization, monitor platform health and cost, and evaluate stakeholder requests
  • Collaborate with engineers, architects, and product managers to drive Core Data platform success; participate in agile/scrum ceremonies
  • Maintain documentation of platform changes, standards, and pipeline configurations to support data quality and governance
Qualifications:
  • 5+ years of data engineering experience developing and operating large‑scale data pipelines
  • Deep hands‑on experience with Databricks and Apache Spark (batch and streaming), including pipeline development in PySpark and/or Scala
  • Strong understanding of Spark architecture—executors, stages, partitioning, shuffle, and performance tuning—with ability to explain tradeoffs to technical and non‑technical stakeholders
  • Proficiency with Databricks platform tooling (API, SDK, CLI) for automation, auditing, governance, and operational troubleshooting
  • Proficient in SQL with advanced performance tuning capabilities
  • Hands‑on production experience with Airflow (MWAA) for orchestrating data pipelines
  • Experience managing Databricks platform governance: ACLs, Unity Catalog, lineage, and access provisioning
  • Proficiency in Python and at least one additional language (Scala, Kotlin, or SQL‑driven pipeline tooling)
  • Experience designing and optimizing scalable ETL/ELT pipelines integrating diverse structured and unstructured data sources
  • AWS‑primary experience (compute, storage, networking, IAM); experience with other cloud providers is transferable
  • Proficiency with Docker and Kubernetes for building and maintaining containerized data platform servicesWorking knowledge of networking concepts to diagnose cross‑platform integration and connectivity issues
  • Familiarity with Snowflake and comparable tooling relative to the Databricks ecosystem
  • Experience designing and implementing CI/CD and DevOps practices (Git‑based workflows)
  • Experience implementing data quality checks, monitoring, and logging for pipeline reliability
  • Self‑starting problem solver with strong analytical and communication skills; willingness to learn new tooling and trends
  • Familiar with Scrum and Agile methodologies
  • Experience with Snowflake is a plus
  • Bachelor’s Degree in Computer Science, Information Systems, or a related field, or equivalent work experience, master’s degree is a plus
Required Skills (Must‑Have)
  • Databricks experience (primary requirement)
  • Advanced SQL skills
  • Experience building and maintaining data pipelines and workflows
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Data Engineer
Data Engineer

Staffingine LLC • Glendale (CA)

On-site
USD 120,000 - 180,000
Senior Data Engineer
Senior Data Engineer

Tata Consultancy Services • Irvine (CA)

On-site
USD 120,000 - 180,000
Databricks Data Engineer
Databricks Data Engineer

Compunnel, Inc. • Spring (TX)

On-site
USD 110,000 - 140,000
Senior Data Engineer
Senior Data Engineer

Acestack • Irvine (CA)

On-site
USD 140,000 - 190,000
Sr. Data Engineer, Data Platform
Sr. Data Engineer, Data Platform

Mirion Technologies • United States

On-site
USD 120,000 - 160,000
Databricks Developer
Databricks Developer

Tredence Inc. • San Jose (CA)

On-site
USD 140,000 - 190,000
Senior Data Engineer
Senior Data Engineer

Siri InfoSolutions Inc • Irvine (CA)

On-site
USD 150,000 - 190,000
Senior Data Engineer
Senior Data Engineer

Motion Recruitment Partners, LLC • Harrisonburg (VA)

On-site
USD 110,000 - 160,000
Bonus eligible
Medical, Dental, and Vision Insurance
Vacation Time
Big Data Specialist
Big Data Specialist

Hyrhub • New York (NY)

On-site
USD 150,000 - 180,000
Databricks Engineer
Databricks Engineer

iLink Digital • Milpitas (CA), Northern (KY)

On-site
USD 120,000 - 160,000