Sr. ML Data Engineer (ETL (Remote)

Cedent

United States

Hybrid

USD 55,104 - 82,656

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Health
Dental
Vision Insurance

Job summary

A leading data engineering company in the United States seeks a Data Engineer to develop and manage ML feature engineering pipelines using Databricks and Apache Spark. The role involves overseeing data integration and optimizing pipelines for both real-time and batch model serving. Candidates should have 7 years of data engineering experience along with proficiency in Python and SQL. The position offers competitive hourly compensation and various health benefits.

Qualifications

  • 7 years in data engineering, with 4 years in ML feature engineering.
  • Experience managing pipelines on Databricks using Apache Spark.
  • Familiarity with ML lifecycle management and MLflow is a plus.

Responsibilities

  • Develop and maintain feature engineering pipelines using Databricks.
  • Integrate diverse data sources to create user behavior profiles.
  • Design and implement ETL, ELT pipelines for medallion architecture.

Skills

Apache Spark
Databricks
Java
Python
Scala
SparkSQL

Job description

Responsibilities
  • Feature Engineering Data Integration Develop and maintain feature engineering pipelines using Data bricks to support ML models effectively
  • Data Pipeline Development Integrate diverse data sources eg clickstreams user behavior demographic data to create user behavior features profiles for complex ML tasks
  • Medallion Architecture Design and implement ETL, ELT pipelines aligned with the bronze silver and gold layers of the medallion architecture
  • Model Support Build data pipelines to support ML model training calibration and deployment leveraging MLflow for experiment tracking and performance monitoring
  • Query Optimization Low Latency Pipelines Design low latency production ready data pipelines to support real-time and batch model inference
  • CICD Practices Apply CICD principles for seamless pipeline deployment
  • Data Governance Ensure pipelines comply with security and regulatory standards particularly for handling PII and maintain metadata and master data across the data catalogue
  • Collaboration Work closely with ml scientists ml engineers and other stakeholders to align data transformation with business objectives
Qualifications
  • 7 years in data engineering and at least 4 years focusing on ML feature engineering ETL pipeline development and data preparation for ML
  • Proven experience managing pipelines on Data bricks using Apache Spark with a strong understanding of the medallion architecture
  • Familiarity with ML lifecycle management with MLflow experience as a strong plus and advanced skills in Apache Spark PySpark for big data processing and analytics
  • Proficient in Python for data manipulation and SQL for query optimization
  • Experience building pipelines for real-time and batch model serving in production environments and knowledge of CICD practices for ETLELT pipeline development
  • Expertise in metadata and master data management within technical data catalogues
  • Understanding of data security and compliance especially with sensitive data like PII
Mandatory Skills
  • Apache Spark
  • Databricks
  • Java
  • Python
  • Scala
  • SparkSQL
Compensation

Hourly Rate Range - $40-$60/ hr

Benefits Offered
  • Health
  • Dental
  • Vision Insurance
Deadline

Applications accepted until10/30/2025 at 11:59 PM CST

We are an Equal Pay Employer. All employment decisions, including compensation, benefits, hiring, training, and promotions, are made based on merit, qualifications, and business needs. We do not discriminate on the basis of gender, race, ethnicity, age, disability, sexual orientation, or any other protected characteristic. We are committed to ensuring equal pay for equal work and regularly review our compensation practices to promote fairness, equity, and transparency across our organization.

Department: Preferred Vendors

This is a contract position

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Machine Learning Engineer
Senior Machine Learning Engineer

Riccione Resources, Inc. • Dallas (TX)

Hybrid
USD 160,000 - 185,000
Sr. Data Engineer (AWS)
Sr. Data Engineer (AWS)

MMD Services • Rosemont (IL)

On-site
USD 120,000 - 150,000
ML Data Engineer
ML Data Engineer

United States Digital Space LLC • United States

Remote
USD 100,000 - 150,000
Sr. Data Engineer
Sr. Data Engineer

Bamboohr17 • Utah

On-site
USD 100,000 - 140,000
Ability to request reasonable accommodations
Equal Opportunity Employer
Senior Data Engineer
Senior Data Engineer

Intellivo • Memphis (TN)

On-site
USD 90,000 - 120,000
Databricks Engineer
Databricks Engineer

CMT Services, Inc. • Adelphi (MD)

On-site
USD 100,000 - 130,000
Senior Data Engineer
Senior Data Engineer

MediData • Woodbridge Township (NJ)

Hybrid
USD 96,000 - 128,000
Senior Data Engineer
Senior Data Engineer

PowerToFly • New Jersey

Hybrid
USD 96,000 - 128,000
Sr. Data Engineer
Sr. Data Engineer

MBO Partners • United States

Remote
USD 90,000 - 130,000
Eligibility for Paid Sick Leave (PSL)
Flexible working hours
AI/ML Data Engineer (Fulltime)
AI/ML Data Engineer (Fulltime)

Aptonet • Tampa (FL)

Hybrid
USD 100,000 - 130,000
Employer-matched 401(k)
Company-paid medical insurance
Company-paid vision insurance
+7