Data Engineer for AI-Driven Data Platform

You.com

San Francisco (CA)

On-site

USD 180,000 - 220,000

Full time

28 hours ago
Be an early applicant
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Benefits offered by this job

Hubs in SF & NYC
Flexible PTO + US holidays
Comprehensive health insurance
12 weeks parental leave (US)
401k with 3% match
Tech stipend for remote work

Job summary

You.com in San Francisco is seeking a hands-on Data Engineer to build and scale a modern data platform. You will partner with Finance, Engineering, Product, and Analytics to develop reliable, high-performance data pipelines and systems.

You’ll work across batch and real-time processing using Databricks, AWS, Kafka, and third-party data, ensuring data quality and availability for analytics, BI, and AI-driven applications.

Qualifications

  • 6+ years of experience in data engineering or a related field.
  • Strong hands-on experience with Databricks, AWS (S3, Glue, Athena, EMR, etc.), and Kafka.
  • Proficiency in Python (PySpark) and SQL for large-scale data processing.
  • Experience building and maintaining ETL/ELT pipelines (DBT/Airflow or similar experience preferred).
  • Experience with data ingestion tools such as Fivetran (or similar).
  • Familiarity with reverse ETL / data activation workflows and syncing data to tools like Salesforce, HubSpot, Braze.
  • Exposure to or experience with AI/ML data pipelines, including RAG architectures, vector databases, or embeddings workflows.
  • Familiarity with agent-based systems, MCP integrations, or LLM-powered applications is a strong plus.
  • Experience working with Finance and building finance specific metrics and pipelines is a strong plus.
  • Understanding of data modeling and working with large-scale datasets (batch and streaming).
  • Experience creating dashboards and supporting reporting workflows (BI tools) for both internal and external audiences.

Responsibilities

  • Build and maintain scalable data pipelines (batch and streaming) using Databricks, Spark, Kafka, and AWS services.
  • Build and maintain pipelines from source systems (Salesforce, billing, product events, API logs) into clean analytics layers.
  • Design, develop, and optimize ETL/ELT workflows using DBT, PySpark, SQL, and tools like Fivetran.
  • Work closely with finance in developing Finance data solutions, Finance metrics and forecasting models.
  • Partner with Finance on revenue accounting, COGS, and margin reporting.
  • Partner closely with marketing and growth teams to enable data use cases such as segmentation, campaign targeting, and lifecycle analytics.
  • Develop and maintain reverse ETL pipelines to sync data from the warehouse to tools like Salesforce, HubSpot, Braze, and other downstream systems.
  • Create and manage curated datasets to support analytics, reporting, and go-to-market initiatives.
  • Build and maintain dashboards and reporting layers to support marketing and business performance tracking.
  • Support AI/ML and agent-based applications by preparing and serving high-quality datasets for MCP integrations and AI driven applications.
  • Monitor pipeline performance, troubleshoot issues, and ensure high data reliability and quality.
  • Implement data quality checks, validations, and alerting mechanisms across both ingestion and activation layers.
  • Collaborate with cross-functional teams to define data contracts and ensure consistency across systems.

Skills

6+ years experience
Data engineering
Problem solving
Communication
Cross-functional
Finance metrics

Tools

Databricks
AWS
Kafka
PySpark
SQL
DBT
Airflow
Fivetran
Salesforce
HubSpot
Braze
MCP integrations

Job description

You.com in San Francisco is seeking a hands-on Data Engineer to build and scale a modern data platform. You will partner with Finance, Engineering, Product, and Analytics to develop reliable, high-performance data pipelines and systems.

You’ll work across batch and real-time processing using Databricks, AWS, Kafka, and third-party data, ensuring data quality and availability for analytics, BI, and AI-driven applications.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Data Engineer - Real-Time Data Platform & AI
Senior Data Engineer - Real-Time Data Platform & AI

youcom • San Francisco (CA)

Hybrid
USD 180,000 - 220,000
In-person gatherings in SF & NYC
Data Engineer
Data Engineer

youcom • San Francisco (CA)

Hybrid
USD 180,000 - 220,000
In-person gatherings in SF & NYC
Data Engineer for AI-Powered Data Platform (Remote)
Data Engineer for AI-Powered Data Platform (Remote)

11x AI Inc. • San Francisco (CA), Northern (KY)

Hybrid
USD 170,000 - 200,000
Data Engineer - Platform & Self-Service Analytics
Data Engineer - Platform & Self-Service Analytics

Baseten • San Francisco (CA)

On-site
USD 180,000 - 250,000
Data Engineer — AI-Powered Data Platform
Data Engineer — AI-Powered Data Platform

Amazon Web Services (AWS) • Arlington (VA)

On-site
USD 132,000 - 179,000
Data Engineer - AI/ML
Data Engineer - AI/ML

Veritas Search Group • Tustin (CA)

On-site
USD 120,000 - 180,000
Data Engineer
Data Engineer

Web3 Tech LLC • Snowflake (AZ)

On-site
USD 100,000 - 130,000
Data Engineer for AI-Powered Data Pipelines & Platform
Data Engineer for AI-Powered Data Pipelines & Platform

Amazon • Hawthorne (CA)

On-site
USD 132,000 - 179,000
Health insurance
RSUs
401(k) matching
+3
Data Engineer
Data Engineer

Prodigy Resources • Denver (CO)

On-site
USD 110,000 - 170,000
Senior Data Engineer - AI-Driven Data Platform
Senior Data Engineer - AI-Driven Data Platform

ITAC Solutions • Birmingham (AL)

On-site
USD 104,000 - 127,000
Next-gen data platform
Source of truth
Azure/Databricks stack
+3