Sr. AI Product Data Engineer

Rpotential

San Francisco (CA)

Hybrid

USD 140,000 - 200,000

Full time

14 days+
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Job summary

rPotential is seeking a Sr. AI Product Data Engineer to own and improve our production data pipelines. You will manage Databricks across jobs, Unity Catalog, and cost while collaborating with product and business teams to understand and resolve data issues.

The role requires 5+ years of data engineering experience, strong Python/SQL skills, and the ability to work in a hybrid setup with three in-person days per week in the SF Bay Area.

Qualifications

  • 5+ years of data engineering experience with production pipeline ownership.
  • Strong Databricks and Unity Catalog experience, including workspace administration.
  • Strong Python and SQL.
  • Experience improving an existing data environment while it remained live.

Responsibilities

  • Own and improve our production data pipelines.
  • Standardize pipeline structure, scheduling, retries, testing, monitoring, and backfills.
  • Build monitoring around freshness, coverage, failures, and data quality.
  • Own Databricks across jobs, Unity Catalog, permissions, environments, and cost.
  • Work in our product monorepo alongside the engineering team.
  • Work with product and business teams to understand new datasets and resolve data issues.
  • Make it easier to take new data sources from prototype to production.

Skills

Data engineering
Python
SQL
Databricks
Unity Catalog
Production pipelines
Data quality
Business judgment
Ambiguous data handling
AI coding tools

Tools

Databricks
Unity Catalog
Postgres

Job description

Sr. AI Product Data Engineer

rPotential | SF Bay Area | 3 days/week in person

About us

rPotential is building a platform that helps large companies understand where to apply AI, how work is changing, and where people can be redeployed to higher-value work. We work directly with Fortune 500 companies and leading AI companies.

We're a small team, move quickly, and expect everyone to be hands-on.

About the role

Our product depends on data. We combine proprietary, third-party, and AI-generated datasets in Databricks and Postgres to power customer-facing products.

We have production pipelines today, but they were built quickly and vary in how they are structured, tested, monitored, and operated. We're looking for a senior data engineer to take ownership of this layer, improve what exists, and establish a consistent way to build and run data pipelines going forward.

This role also requires understanding the business context behind the data. You'll need to understand what the data represents, where it came from, where it can be misleading, and how it is used in the product.

What you'll do
  • Own and improve our production data pipelines.
  • Standardize pipeline structure, scheduling, retries, testing, monitoring, and backfills.
  • Build monitoring around freshness, coverage, failures, and data quality.
  • Own Databricks across jobs, Unity Catalog, permissions, environments, and cost.
  • Work in our product monorepo alongside the engineering team.
  • Work with product and business teams to understand new datasets and resolve data issues.
  • Make it easier to take new data sources from prototype to production.
What we're looking for
  • 5+ years of data engineering experience with production pipeline ownership.
  • Strong DataBricks and Unity Catalog experience, including workspace administration.
  • Strong Python and SQL.
  • Experience improving an existing data environment while it remained live.
  • Strong business judgment around data and an interest in understanding what the data actually means.
  • Comfortable working through ambiguous or messy data with product and business teams rather than waiting for a finished spec.
  • Comfortable with AI coding tools
  • Bay Area based or commutable and comfortable working in person three days per week.
Nice to ave
  • Azure, terraform, Postgres
  • Entity resolution or identity matching.
  • Labor-market, company, people, or other large third-party datasets.
  • Data systems supporting AI or LLM products.
What we offer

High ownership, direct work with the founder and engineering team, and the opportunity to define how data engineering is done at rPotential as the team grows.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior AI Data Engineer: Own Pipelines & Data Quality
Senior AI Data Engineer: Own Pipelines & Data Quality

Rpotential • San Francisco (CA)

Hybrid
USD 140,000 - 200,000
Sr. Data Engineer
Sr. Data Engineer

Echo Global Logistics • United States

Remote
USD 120,000 - 180,000
Bonus eligibility
Senior Artificial Intelligence Data Engineer
Senior Artificial Intelligence Data Engineer

Vizient • Irving (TX)

On-site
USD 102,400 - 179,000
Data Engineer
Data Engineer

youcom • San Francisco (CA)

On-site
USD 180,000 - 220,000
In-person gatherings in SF & NYC
Staff Application Engineer - Enterprise Data & AI (Intelligence Platform)
Staff Application Engineer - Enterprise Data & AI (Intelligence Platform)

Riot Games • Los Angeles (CA)

On-site
USD 180,000 - 260,000
Senior AI Engineer
Senior AI Engineer

Tenth Revolution Group • Chicago (IL)

On-site
USD 140,000 - 200,000
Senior Data Infrastructure Engineer
Senior Data Infrastructure Engineer

Rise Technical • United States

On-site
USD 140,000 - 180,000
Equity
401K
PTO
+1
Senior Data Engineer
Senior Data Engineer

SDL Search Partners • Boston (MA)

On-site
USD 130,000 - 185,000
Competitive compensation
AI Data Architect[REQ_22]
AI Data Architect[REQ_22]

3Pillar Global, Inc. • Northern (KY)

On-site
USD 150,000 - 190,000
Medical Insurance
Dental Insurance
Vision Insurance
+7
Staff Scientific Data Engineer
Staff Scientific Data Engineer

Insilico Search Partners • Boston (MA)

On-site
USD 140,000 - 200,000