Senior Data Engineer-5

REALIGN LLC

Irvine (CA)

On-site

USD 140,000 - 210,000

Full time

1 hour ago
Be an early applicant
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Job summary

REALIGN LLC is seeking a hands-on Senior Data Engineer to design, build, and optimize enterprise-scale data solutions on the Databricks Lakehouse Platform. You will develop reliable batch and streaming pipelines and modernize legacy ETL workloads for analytics, reporting, risk, regulatory, and investment-management use cases.

The ideal candidate brings deep Databricks, Spark, PySpark, SQL, Python, dbt, Airflow, Delta Lake, Unity Catalog, cloud storage, data quality, CI/CD, and production support

Qualifications

  • Typically 7–10 years of data engineering, data warehousing, or distributed data‑processing experience.
  • Strong hands‑on experience with Databricks, Apache Spark, PySpark, Delta Lake, Python, and advanced SQL.
  • Experience building production‑grade ETL and ELT pipelines for large datasets.
  • Experience with dbt Core or dbt Cloud, including models, macros, tests, documentation, and incremental processing.
  • Experience with Apache Airflow, Astronomer, Databricks Workflows, or comparable orchestration platforms.
  • Experience with Unity Catalog, data lineage, role‑based access, and data‑governance controls.
  • Experience with cloud data services on AWS, Azure, or Google Cloud.
  • Working knowledge of Git, CI/CD, automated testing, monitoring, and production‑support practices.
  • Strong troubleshooting, communication, collaboration, and technical‑documentation skills.

Responsibilities

  • Design, build, test, deploy, and maintain scalable ETL and ELT pipelines using Databricks, PySpark, Spark SQL, Python, and SQL.
  • Develop reusable ingestion and transformation frameworks for structured, semi-structured, and streaming data.
  • Implement batch, incremental, change-data-capture, and streaming processing patterns.
  • Build and maintain Delta Lake tables using medallion architecture across Bronze, Silver, and Gold layers.
  • Develop dbt models, tests, macros, packages, documentation, and incremental processing patterns.
  • Create and maintain Apache Airflow DAGs and Databricks Workflows with dependency management, retries, alerting, and operational controls.
  • Collaborate to integrate data from APIs, databases, files, event streams, and cloud services.
  • Provide technical designs, mapping specs, lineage docs, deployment instructions, and runbooks.

Skills

Databricks
Apache Spark
PySpark
SQL
Python
dbt
Apache Airflow
Delta Lake
Unity Catalog
Cloud storage
Data quality
CI/CD
Production support

Tools

Databricks
dbt Core
dbt Cloud
Apache Airflow
Astronomer
Databricks Workflows
Terraform

Job description

We are seeking a hands-on Senior Data Engineer to design, build, optimize, and support enterprise-scale data solutions on the Databricks Lakehouse Platform. The role will develop reliable batch and streaming pipelines, modernize legacy ETL workloads, implement governed data models, and deliver trusted data products for analytics, reporting, risk, regulatory, and investment-management use cases.

The ideal candidate has deep experience with Databricks, Apache Spark, PySpark, SQL, Python, dbt, Apache Airflow, Delta Lake, Unity Catalog, cloud storage, data quality, CI/CD, and production support.

Key Responsibilities
Data Engineering and Development
  • Design, build, test, deploy, and maintain scalable ETL and ELT pipelines using Databricks, PySpark, Spark SQL, Python, and SQL.
  • Develop reusable ingestion and transformation frameworks for structured, semi-structured, and streaming data.
  • Implement batch, incremental, change-data-capture, and streaming processing patterns.
  • Build and maintain Delta Lake tables using medallion architecture across Bronze, Silver, and Gold layers.
  • Develop dbt models, tests, macros, packages, documentation, and incremental processing patterns.
  • Create and maintain Apache Airflow DAGs and Databricks Workflows with dependency management, retries, alerting, and operational controls.
  • Integrate data from APIs, databases, files, event streams, and cloud data services.
  • Produce technical designs, mapping specifications, lineage documentation, deployment instructions, and operational runbooks.
Performance, Reliability, and Data Quality
  • Tune Spark workloads, joins, partitioning, file sizes, caching, cluster configurations, and query plans.
  • Apply Delta Lake optimization techniques, including compaction, data skipping, clustering, retention, and vacuum controls.
  • Implement automated data quality, reconciliation, schema validation, observability, and freshness checks.
  • Monitor pipeline health and resolve failures, performance degradation, data defects, and service-level breaches.
  • Perform root‑cause analysis and implement durable preventive measures.
  • Support release readiness, production cutover, incident resolution, and ongoing platform operations.
  • Improve compute utilization and cost efficiency across batch and streaming workloads.
Governance, Security, and Delivery Practices
  • Apply Unity Catalog standards for catalogs, schemas, tables, views, lineage, classification, and controlled access.
  • Implement secure handling of credentials, secrets, personally identifiable information, and regulated data.
  • Contribute to CI/CD pipelines, automated testing, code‑quality checks, and environment promotion.
  • Use Git‑based development, peer reviews, branching standards, and release‑management practices.
  • Collaborate with platform engineers to deploy data assets through Terraform and Databricks Asset Bundles where applicable.
  • Follow enterprise architecture, security, data‑governance, and regulatory requirements.
Collaboration and Mentoring
  • Partner with architects, product owners, analysts, data scientists, governance teams, and business stakeholders.
  • Translate business requirements into scalable data models, pipelines, and technical work packages.
  • Conduct code reviews and enforce engineering, documentation, testing, and support standards.
  • Mentor junior and mid‑level engineers and share reusable patterns and best practices.
  • Communicate delivery status, risks, dependencies, and technical trade‑offs clearly.
Required Qualifications
  • Typically 7–10 years of data engineering, data warehousing, or distributed data‑processing experience.
  • Strong hands‑on experience with Databricks, Apache Spark, PySpark, Delta Lake, Python, and advanced SQL.
  • Experience building production‑grade ETL and ELT pipelines for large datasets.
  • Experience with dbt Core or dbt Cloud, including models, macros, tests, documentation, and incremental processing.
  • Experience with Apache Airflow, Astronomer, Databricks Workflows, or comparable orchestration platforms.
  • Experience with Unity Catalog, data lineage, role‑based access, and data‑governance controls.
  • Experience with cloud data services on AWS, Azure, or Google Cloud.
  • Working knowledge of Git, CI/CD, automated testing, monitoring, and production‑support practices.
  • Strong troubleshooting, communication, collaboration, and technical‑documentation skills.
Preferred Qualifications
  • Experience in banking, financial services, insurance, asset management, risk, compliance, or another regulated industry.
  • Experience modernizing Hadoop, legacy data warehouses, or traditional ETL platforms.
  • Experience with Kafka, Structured Streaming, Auto Loader, Delta Live Tables, or Lakeflow Declarative Pipelines.
  • Familiarity with Terraform, Databricks Asset Bundles, cloud networking, IAM, secrets management, and infrastructure automation.
  • Databricks Data Engineer Associate or Professional certification.
  • Experience delivering data reconciliation, regulatory reporting, test automation, and audit‑ready controls.
Job Type: Full Time
Job Category: IT
Job Description

Role: Senior Data Engineer

Location: Irvine, CA (Onsite)

Fulltime – Permanent

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Data Engineer
Senior Data Engineer

Siri InfoSolutions Inc • Irvine (CA)

On-site
USD 150,000 - 190,000
Senior Data Engineer
Senior Data Engineer

Acestack • Irvine (CA)

On-site
USD 140,000 - 190,000
Senior Data Engineer
Senior Data Engineer

Tata Consultancy Services • Irvine (CA)

On-site
USD 120,000 - 180,000
Senior Data Engineer
Senior Data Engineer

Prosum • Glendale (CA)

On-site
USD 150,000 - 210,000
Databricks Data Engineer
Databricks Data Engineer

Compunnel, Inc. • Spring (TX)

On-site
USD 110,000 - 140,000
Data Team Lead
Data Team Lead

Incedo Inc. • San Rafael (CA)

Hybrid
USD 130,000 - 150,000
Medical insurance
Vision insurance
401(k)
Databricks Developer
Databricks Developer

Tredence Inc. • San Jose (CA)

On-site
USD 140,000 - 190,000
Senior Data Engineer
Senior Data Engineer

Motion Recruitment Partners, LLC • Harrisonburg (VA)

On-site
USD 110,000 - 160,000
Bonus eligible
Medical, Dental, and Vision Insurance
Vacation Time
Full-Stack Engineer Senior
Full-Stack Engineer Senior

Donnelly- • New York (NY)

On-site
USD 140,000 - 200,000
Senior Data Engineer - Databricks
Senior Data Engineer - Databricks

DATAECONOMY Inc • Northern (KY)

On-site
USD 140,000 - 180,000