Platform Data Engineer for AI Pipelines

Socket.dev

San Francisco (CA)

On-site

USD 127,000 - 190,000

Full time

10 days ago

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Health benefits
Equity/Stock options
Flexible time off
Paid holidays

Job summary

Invoca is seeking a data engineering leader to own the core data infrastructure powering AI capabilities. You’ll build scalable data pipelines, extend the lakehouse architecture on Databricks, and curate datasets used by data science and ML teams. The role focuses on data quality, governance, and reliability at scale within a remote-friendly U.S.

organization. You’ll mentor teams, drive CI/CD for data code, and partner with leadership to prioritize platform investments.

Qualifications

  • 5+ years of professional experience in data engineering or related roles with ownership of production data systems.
  • Proficient in Python for data engineering, PySpark, and data processing libraries.
  • Strong SQL skills, performance tuning, and data modeling for relational and lakehouse environments.
  • Experience with Databricks or equivalent lakehouse stack; familiar with Spark at scale.
  • Understanding of data governance, privacy, and responsible data handling.
  • Experience with streaming technologies (Kafka, Kinesis, Spark Structured Streaming).
  • Working knowledge of AWS and security fundamentals.

Responsibilities

  • Design, build, and operate scalable batch and streaming data pipelines across the Invoca Platform.
  • Evolve the lakehouse architecture using Databricks and related tools as a shared platform.
  • Build curated training and evaluation datasets and feature pipelines used by data science and ML teams.
  • champion data quality, CI/CD for data code, data schema management, and dataset versioning.
  • Ensure responsible data handling with governance, access control, and PI protection.

Skills

Python for data engineering
PySpark
SQL proficiency
Data modeling
CI/CD for data code
Data governance
Data quality monitoring
Streaming data tech
Cloud basics (AWS)

Education

Bachelor's degree or equivalent experience

Tools

Databricks (Delta Lake)
Spark
Airflow
Dagster
Kinesis
S3

Job description

Invoca is seeking a data engineering leader to own the core data infrastructure powering AI capabilities. You’ll build scalable data pipelines, extend the lakehouse architecture on Databricks, and curate datasets used by data science and ML teams. The role focuses on data quality, governance, and reliability at scale within a remote-friendly U.S.

organization. You’ll mentor teams, drive CI/CD for data code, and partner with leadership to prioritize platform investments.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Data Platform Engineering Lead: Build & Scale Pipelines
Data Platform Engineering Lead: Build & Scale Pipelines

Reflection • San Francisco (CA)

On-site
USD 180,000 - 240,000
Stock options
Health & wellness
Meals in office
+2
Senior Data Platform Engineer - Remote & Scalable Pipelines
Senior Data Platform Engineer - Remote & Scalable Pipelines

Parafin, Inc. • San Francisco (CA)

On-site
USD 220,000 - 265,000
Work from home flexibility
Unlimited PTO
Free lunches
+4
Senior Data Platform Engineer: Scale AI Data Pipelines
Senior Data Platform Engineer: Scale AI Data Pipelines

Weights & Biases • New York (NY)

On-site
USD 165,000 - 242,000
Medical, dental, and vision insurance
Company-paid Life Insurance
401(k) with employer match
+5
Staff Data Engineering Lead - Scalable Pipelines
Staff Data Engineering Lead - Scalable Pipelines

EngineersOfAI • Seattle (WA)

On-site
USD 130,000 - 170,000
Senior Data Platform Engineer — Scale AI Data Pipelines
Senior Data Platform Engineer — Scale AI Data Pipelines

Weights & Biases • San Francisco (CA)

On-site
USD 165,000 - 242,000
Medical, dental, and vision insurance
401(k) with employer match
Paid Parental Leave
+4
Senior AI Data Engineer: Scalable Pipelines & Infra
Senior AI Data Engineer: Scalable Pipelines & Infra

IBM • Los Angeles (CA)

On-site
USD 150,000 - 190,000
Senior Data Platform Engineer - AI-Driven Pipelines
Senior Data Platform Engineer - AI-Driven Pipelines

Craft • United States

Hybrid
USD 150,000 - 190,000
Equity
Unlimited vacation
Health + dental + vision insurance (99
+1
Senior AI-Ready Data Engineer – Pipelines & Lakehouse
Senior AI-Ready Data Engineer – Pipelines & Lakehouse

Stryker Corporation • Charlotte (NC), Northern (KY)

Hybrid
Confidential
Paid time off
Benefits
Senior Data Engineer: AI Pipelines & Data Governance
Senior Data Engineer: AI Pipelines & Data Governance

IBM • Austin (TX)

On-site
USD 120,000 - 160,000
Remote-friendly
Competitive salary
IBM culture
Remote Data Engineer: AI Pipelines & AWS
Remote Data Engineer: AI Pipelines & AWS

Appsierra Group • United States

On-site
USD 140,000 - 180,000
Equity
Performance bonuses
Health insurance reimbursement
+3