Data Engineer – Remote Databricks Pipelines

Irth Solutions

United States

Remote

USD 45,000 - 61,000

Full time

14 days+
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Benefits offered by this job

Competitive salary
Medical, dental, vision insurance
401(k) plan with company match
Generous PTO
Company-paid holidays
Flexible work options
On-call compensation

Job summary

Irth Solutions is seeking a mid-level Data Engineer to design, build, and maintain data ingestion and processing pipelines in Databricks, transforming external sources into clean, analysis-ready inputs for downstream intelligence. You will work primarily on the Stakeholder Engagement data and collaborate with Data Scientist and application teams.

The role emphasizes medallion architecture, Delta Lake optimization, reliability, and ongoing collaboration with cross-functional teams to deliver

Qualifications

  • 3 to 5 years of experience in data engineering, with solid experience building and operating production-grade data pipelines.
  • Familiarity with data modeling, data quality, and schema evolution.
  • Strong proficiency in Python and SQL
  • Hands-on experience with Databricks (Spark, Delta Lake)
  • Experience with at least one major cloud (Azure preferred; AWS/GCP beneficial)
  • Experience integrating with external APIs at scale: authentication, pagination, rate limiting, retries, error handling.
  • Databricks certification preferred (Data Engineer Associate)

Responsibilities

  • Data Pipeline Development (Primary Responsibility)
  • Design, build, and maintain ingestion pipelines from high-volume external APIs, capable of running continuously and reliably at scale.
  • Implement ingestion and transformation workflows using Databricks (Spark/PySpark, SQL, Delta Live Tables), applying medallion architecture patterns (Bronze → Silver → Gold) to move from raw ingested content to clean, structured, analysis-ready data.
  • Build the infrastructure for deduplication and relevance filtering of incoming content, implementing filtering logic and quality criteria defined in collaboration with the Data Scientist.
  • Implement schema evolution handling and data validation rules as data sources and formats change over time.
  • Platform & Storage Implementation
  • Configure and manage Delta Lake storage structures, tables, partitions, and optimization routines (OPTIMIZE, Z-ORDER, VACUUM).
  • Design and evolve data schemas that balance query performance, cost, and maintainability as data volume grows.
  • Maintain clear metadata and documentation of table structures to support easy consumption by the Data Science and application teams.
  • Reliability, Monitoring & Operational Support
  • Ensure pipeline reliability and observability: error handling, retries, monitoring, and alerting for a continuously running system.
  • Adapt pipelines to evolving external API contracts, rate limits, authentication changes, and new data sources.
  • Troubleshoot pipeline failures, perform recovery, and tune performance as needed.
  • Orchestration & Automation
  • Build, schedule, and monitor workflows using Databricks Workflows, Delta Live Tables, or similar orchestration tools.
  • Contribute to CI/CD pipelines for code deployment, versioning, and environment management.
  • Collaboration & Documentation
  • Work closely with the Data Scientist to expose clean, well-structured data feeding LLM/NLP pipelines and downstream models.
  • Participate in technical decisions around data architecture and propose structuring solutions as the team’s needs evolve.
  • Document pipelines, data dictionaries, job schedules, and transformation logic.
  • Support the onboarding of new data sources and pipelines as the product expands to additional solution areas.

Skills

Databricks
Python
SQL
Delta Lake
Spark
Data pipelines
CI/CD
Databricks Workflows

Tools

Delta Live Tables

Job description

Irth Solutions is seeking a mid-level Data Engineer to design, build, and maintain data ingestion and processing pipelines in Databricks, transforming external sources into clean, analysis-ready inputs for downstream intelligence. You will work primarily on the Stakeholder Engagement data and collaborate with Data Scientist and application teams.

The role emphasizes medallion architecture, Delta Lake optimization, reliability, and ongoing collaboration with cross-functional teams to deliver

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Data Engineer, Databricks Pipelines — Remote
Data Engineer, Databricks Pipelines — Remote

Irth Solutions • United States

Remote
USD 46,000 - 62,000
Competitive Salary
Medical, Dental, and Vision Insurance
401(k) Plan with Company Match
+2
Data Engineer – Insights (AI/ML) | Remote
Data Engineer – Insights (AI/ML) | Remote

Irth Solutions • United States

Remote
USD 120,000 - 180,000
Medical Insurance
Flexible Work Options
401(k) Plan
+1
Remote Databricks Engineer: Scalable Data Pipelines
Remote Databricks Engineer: Scalable Data Pipelines

General Dynamics Corporation • Silver Spring (MD), Northern (KY)

Hybrid
USD 140,000 - 190,000
Medical plan options
Paid time off
Disability insurance
Data Engineer (Databricks) - Remote, AI-Powered Pipelines
Data Engineer (Databricks) - Remote, AI-Powered Pipelines

Full Tilt Data, LLC • United States

On-site
USD 110,000 - 160,000
Health benefits
Discretionary bonuses
Reimbursement for professional develop
Databricks Solutions Architect – Remote Data Pipelines
Databricks Solutions Architect – Remote Data Pipelines

United States Digital Space LLC • United States

Remote
USD 240,000 - 260,000
Competitive hourly rate
Flexible contract or full-time work
Work with leading enterprise customers
Databricks Data Engineer — Pipelines & Big Data
Databricks Data Engineer — Pipelines & Big Data

Cyber Space Technologies LLC • Columbus (OH)

On-site
USD 90,000 - 115,000
Databricks Data Engineer - Spark, ETL & Cloud Pipelines
Databricks Data Engineer - Spark, ETL & Cloud Pipelines

Smart IT Frame LLC • Reston (VA)

On-site
USD 90,000 - 120,000
IoT Data Engineer: Databricks & Real-Time Pipelines Remote
IoT Data Engineer: Databricks & Real-Time Pipelines Remote

Coherent Solutions, Inc. • Oak Brook (IL)

Hybrid
USD 110,000 - 170,000
Health insurance
Flexible work options
Data Engineer - Databricks & Cloud Data Pipelines
Data Engineer - Databricks & Cloud Data Pipelines

ContractStaffingRecruiters.com • Hartford (CT)

Remote
USD 110,000 - 165,000
Data Engineer I: AI-Driven Pipelines on Databricks & AWS
Data Engineer I: AI-Driven Pipelines on Databricks & AWS

Travelers • Hartford (CT)

On-site
USD 109,000 - 180,000
Health Insurance
401(k) Matching
Paid Time Off
+2