Data Engineer

Staffingine LLC

Glendale (CA)

On-site

USD 120,000 - 180,000

Full time

7 days ago
Be an early applicant
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Job summary

Staffingine LLC in Glendale, CA is seeking an experienced Data Engineer to design, build, and deploy robust data pipelines using PySpark, Scala, SQL, Python, and Databricks, with AWS and Kubernetes considerations. You will collaborate with engineers, product managers, and stakeholders, ensure data quality and governance, and participate in agile ceremonies to deliver scalable data platform solutions.

The role emphasizes hands-on pipeline development, platform administration, and cross-team

Qualifications

  • 5+ years of data engineering experience building large-scale pipelines.
  • Deep hands-on Databricks and Spark (batch and streaming).
  • Strong SQL with performance tuning and Python proficiency.
  • Experience with Airflow MWAA and containerized deployments.
  • AWS-focused data platform knowledge and governance concepts.
  • Knowledge of Kubernetes, Docker, and data security/compliance.

Responsibilities

  • Design, write, test, and deploy data pipelines using PySpark, Scala, SQL, Python.
  • Collaborate with stakeholders to translate requirements into scalable solutions.
  • Governance: Unity Catalog, ACLs, data discovery, lineage, privacy tooling.
  • Maintain platform health, cost, and performance; monitor pipelines and logs.
  • Support CI/CD and DevOps practices for data workflows.
  • Explain Spark architecture to non-technical stakeholders.

Skills

Databricks experience
Advanced SQL
Python
Spark / PySpark
Scala
Data pipelines
Agile / Scrum
Communication

Education

Bachelor's Degree in CS/IS or related
Master's Degree a plus

Tools

Databricks
Apache Spark
Airflow
AWS
Kubernetes
Docker
Unity Catalog
Snowflake
MWAA
CI/CD tooling

Job description

Job Title: Data Engineer
Job Location: Glendale, CA
Job Type: Full-Time

Job Description:


  • Databricks experience (primary requirement)

  • Apache Airflow

  • Advanced SQL skills

  • Python

  • Spark / PySpark

  • Scala

  • Experience building and maintaining data pipelines and workflows


Interview Process:


  • Expected to follow the team's recent hiring process for similar data engineering positions.

  • Typically consists of:

  • Initial screening/vetting round

  • Technical interview

  • Final interview (potentially onsite)

  • Total process generally spans 2-3 interview rounds.


Key Responsibilities:


  • Design, write, test, and deploy data pipelines using PySpark, Scala, SQL, Python

  • Meet with stakeholders to gather requirements and translate them into scalable data platform solutions

  • Understanding of Databricks platform and developer tooling to diagnose errors, audit platform activity, and automate updates across pipelines, objects, and integrations

  • Ability to explain Spark architecture and pipeline behavior to stakeholders to diagnose root causes and recommend solutions

  • Provide solution architecture across AWS, Databricks, Kubernetes, and Airflow (MWAA), including cross-platform integrations

  • Manage Databricks platform governance, including Unity Catalog, ACLs, lineage, and data discovery and privacy tooling

  • Build and maintain Kubernetes containers and containerized utilities supporting deployed data platform services

  • Apply networking knowledge to troubleshoot connectivity and integration errors across platform components

  • Perform platform administration: provision and remove access, assess resource utilization, monitor platform health and cost, and evaluate stakeholder requests

  • Collaborate with engineers, architects, and product managers to drive Core Data platform success; participate in agile/scrum ceremonies

  • Maintain documentation of platform changes, standards, and pipeline configurations to support data quality and governance


Qualifications:


  • 5+ years of data engineering experience developing and operating large-scale data pipelines

  • Deep hands-on experience with Databricks and Apache Spark (batch and streaming), including pipeline development in PySpark and/or Scala

  • Strong understanding of Spark architecture-executors, stages, partitioning, shuffle, and performance tuning-with ability to explain tradeoffs to technical and non-technical stakeholders

  • Proficiency with Databricks platform tooling (API, SDK, CLI) for automation, auditing, governance, and operational troubleshooting

  • Proficient in SQL with advanced performance tuning capabilities

  • Hands-on production experience with Airflow (MWAA) for orchestrating data pipelines

  • Experience managing Databricks platform governance: ACLs, Unity Catalog, lineage, and access provisioning

  • Proficiency in Python and at least one additional language (Scala, Kotlin, or SQL-driven pipeline tooling)

  • Experience designing and optimizing scalable ETL/ELT pipelines integrating diverse structured and unstructured data sources

  • AWS-primary experience (compute, storage, networking, IAM); experience with other cloud providers is transferable

  • Proficiency with Docker and Kubernetes for building and maintaining containerized data platform services

  • Working knowledge of networking concepts to diagnose cross-platform integration and connectivity issues

  • Familiarity with Snowflake and comparable tooling relative to the Databricks ecosystem

  • Experience designing and implementing CI/CD and DevOps practices (Git-based workflows)

  • Experience implementing data quality checks, monitoring, and logging for pipeline reliability

  • Self-starting problem solver with strong analytical and communication skills; willingness to learn new tooling and trends

  • Familiar with Scrum and Agile methodologies

  • Experience with Snowflake is a plus

  • Bachelor's Degree in Computer Science, Information Systems, or a related field, or equivalent work experience, Master's Degree is a plus

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Databricks Engineer
Databricks Engineer

iLink Digital • Milpitas (CA), Northern (KY)

On-site
USD 120,000 - 160,000
Databricks Data Engineer
Databricks Data Engineer

Henderson Scott • Irving (TX)

On-site
USD 100,000 - 130,000
Databricks Data Engineer
Databricks Data Engineer

Veriipro • United States

On-site
USD 120,000 - 170,000
Senior Data Engineer
Senior Data Engineer

Tata Consultancy Services • Irvine (CA)

On-site
USD 120,000 - 180,000
Databricks Data Engineer
Databricks Data Engineer

Compunnel, Inc. • Spring (TX)

On-site
USD 110,000 - 140,000
Databricks Technical Lead
Databricks Technical Lead

Anblicks • Dallas (TX)

On-site
USD 120,000 - 160,000
Databricks Architect
Databricks Architect

Anblicks Inc. • Dallas (TX)

On-site
USD 120,000 - 160,000
Data Engineering Manager
Data Engineering Manager

Burtch Works • Chicago (IL)

On-site
USD 150,000 - 190,000
Databricks Developer
Databricks Developer

Tredence Inc. • San Jose (CA)

On-site
USD 140,000 - 190,000
BI Data Engineer II
BI Data Engineer II

Jobtailor • Boston (MA)

On-site
USD 120,000 - 180,000