Data Engineer New San Francisco (Hybrid)

YOU.com

San Francisco, Northern (CA, KY)

Hybrid

USD 180,000 - 220,000

Full time

2 days ago
Be an early applicant
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Benefits offered by this job

Hubs in San Francisco and New York
Flexible PTO with US holidays
Health insurance
401k 3% match
Remote work stipend
Tech stipend
Wellness allowance
AI collaboration opportunities

Job summary

You.com is seeking a hands-on Data Engineer to build and scale a modern data platform in a collaborative, cross‑functional environment.

You will develop reliable batch and real‑time data pipelines (Databricks, AWS, Kafka), enable data activation for business tools, and support AI/ML data workflows while ensuring data quality and accessibility across the organization.

Qualifications

  • 6+ years of experience in data engineering or a related field.
  • Strong hands-on experience with Databricks, AWS (S3, Glue, Athena, EMR, etc.), and Kafka.
  • Proficiency in Python (PySpark) and SQL for large-scale data processing.
  • Experience building and maintaining ETL/ELT pipelines (DBT/Airflow or similar).
  • Experience with data ingestion tools such as Fivetran (or similar).
  • Familiarity with reverse ETL / data activation workflows and syncing data to Salesforce, HubSpot, Braze.
  • Exposure to AI/ML data pipelines, including RAG architectures, vector databases, or embeddings workflows.
  • Familiarity with agent-based systems, MCP integrations, or LLM-powered applications is a strong plus.
  • Experience working with Finance and building finance specific metrics and pipelines is a strong plus.
  • Understanding of data modeling and working with large-scale datasets (batch and streaming).
  • Experience creating dashboards and supporting reporting workflows (BI tools).
  • Strong problem-solving skills and ability to debug production data issues.
  • Strong communication skills and ability to work collaboratively across teams.

Responsibilities

  • Build and maintain scalable data pipelines (batch and streaming) using tools such as Databricks, Spark, Kafka, and AWS services.
  • Build and maintain pipelines from source systems (Salesforce, billing, product events, API logs) into clean analytics layers.
  • Design, develop, and optimize ETL/ELT workflows using DBT, PySpark, SQL, and tools like Fivetran.
  • Work closely with finance in developing Finance data solutions, Finance metrics and forecasting models.
  • Partner with Finance on revenue accounting, COGS, and margin reporting.
  • Partner closely with marketing and growth teams to enable data use cases such as segmentation, campaign targeting, and lifecycle analytics.
  • Develop and maintain reverse ETL pipelines to sync data from the warehouse to tools like Salesforce, HubSpot, Braze, and other downstream systems.
  • Create and manage curated datasets to support analytics, reporting, and go-to-market initiatives.
  • Build and maintain dashboards and reporting layers to support marketing and business performance tracking.
  • Support AI/ML and agent-based applications by preparing and serving high-quality datasets for MCP integrations and AI driven applications.
  • Monitor pipeline performance, troubleshoot issues, and ensure high data reliability and quality.
  • Implement data quality checks, validations, and alerting mechanisms across both ingestion and activation layers.
  • Collaborate with cross-functional teams to define data contracts and ensure consistency across systems.

Skills

Databricks
AWS
Kafka
Python (PySpark)
SQL
ETL/ELT pipelines
Fivetran
Reverse ETL
Finance data
BI dashboards

Tools

DBT
Airflow
Fivetran
Salesforce
HubSpot
Braze
Spark

Job description

At You.com, we are building the AI Search Infrastructure that powers modern AI systems. Our goal is to create the trusted knowledge layer that agents, applications, and enterprises rely on to retrieve real-time, accurate, and citation-backed information.

Our platform combines proprietary vertical indexes with LLM-optimized retrieval systems to power AI agents, applications, and enterprise workflows. We are solving hard problems across search, large language models, and large‑scale infrastructure to make AI systems more reliable, transparent, and useful.

Our team includes engineers, researchers, product builders, and operators who care about solving meaningful problems and delivering real‑world impact. Whether you are improving core infrastructure, shaping product experiences, or helping bring new AI capabilities to market, your work will help define how modern AI finds and uses knowledge.

About the Role

We are looking for a hands‑on Data Engineer to help build and scale our modern data platform. In this role, you will work closely with Finance, Engineering, Product, and Analytics teams to develop reliable, high‑performance data pipelines and systems.

You’ll contribute to both batch and real‑time data processing using technologies like Databricks, AWS, kafka and several 3rd party data, while helping ensure data quality, accessibility, and usability across the organization. You’ll play a key role in enabling data activation, ensuring that high‑quality data flows not only into the warehouse but also outward to business tools such as Salesforce etc. Additionally, you will help power next‑generation AI‑driven applications, including agent‑based systems and AI driven tools using OSS tech, by building robust data foundations and pipelines. This is a great opportunity for someone who enjoys solving data challenges end‑to‑end from ingestion to insights.

Responsibilities

  • Build and maintain scalable data pipelines (batch and streaming) using tools such as Databricks, Spark, Kafka, and AWS services
  • Build and maintain pipelines from source systems (Salesforce, billing, product events, API logs) into clean analytics layers
  • Design, develop, and optimize ETL/ELT workflows usingDBT, PySpark, SQL, and tools like Fivetran
  • Work closely withfinance in developing Finance data solutions, Finance metrics and forecasting models
  • Partner with Finance on revenue accounting, COGS, and margin reporting
  • Partner closely with marketing and growth teams to enable data use cases such as segmentation, campaign targeting, and lifecycle analytics
  • Develop and maintainreverse ETL pipelines to sync data from the warehouse to tools like Salesforce, HubSpot, Braze, and other downstream systems
  • Create and manage curated datasets to support analytics, reporting, and go‑to‑market initiatives
  • Build and maintain dashboards and reporting layers to support marketing and business performance tracking
  • Support AI/ML and agent‑based applications by preparing and serving high‑quality datasets for MCP (Model Context Protocol) integrations and AI driven applications
  • Monitor pipeline performance, troubleshoot issues, and ensure high data reliability and quality
  • Implement data quality checks, validations, and alerting mechanisms across both ingestion and activation layers
  • Collaborate with cross‑functional teams to define data contracts and ensure consistency across systems
Qualifications
  • 6+ years of experience in data engineering or a related field
  • Strong hands‑on experience with Databricks, AWS (S3, Glue, Athena, EMR, etc.), and Kafka
  • Proficiency in Python (PySpark) and SQL for large‑scale data processing
  • Experience building and maintaining ETL/ELT pipelines (DBT/Airflow or similar experience preferred)
  • Experience with data ingestion tools such as Fivetran (or similar)
  • Familiarity with reverse ETL / data activation workflows and syncing data to tools like Salesforce, HubSpot, Braze
  • Exposure to or experience with AI/ML data pipelines, including RAG architectures, vector databases, or embeddings workflows
  • Familiarity with agent‑based systems, MCP integrations, or LLM‑powered applications is a strong plus
  • Experience working with Finance and building finance specific metrics and pipelines is a strong plus
  • Understanding of data modeling and working with large‑scale datasets (batch and streaming)
  • Experience creating dashboards and supporting reporting workflows (BI tools) for both internal and external audiences
  • Strong problem‑solving skills and ability to debug production data issues
  • Strong communication skills and ability to work collaboratively across teams

Our salary bands are structured based on a combination of geographic tiers and internal leveling. Compensation is determined by multiple factors assessed during the interview process, with the final offer reflecting these considerations.

Salary Band

$180,000 - $220,000 USD

Company Perks:

Hubs in San Francisco and New York City offering regular in‑person gatherings and co‑working sessions

Flexible PTO with U.S. holidays observed and a week shutdown in December to rest and recharge*

A competitive health insurance plan covers 100% of the policyholder and 75% for dependents*

12 weeks of paid parental leave in the US*

401k program, 3% match - vested immediately!*

$500 work‑from‑home stipend to be used up to a year of your start date*

  • $600 technology stipend to support a portion of our hybrid/remote team's cell phone and internet expenses*

$1,200 per year Health & Wellness Allowance to support your personal goals*

The chance to collaborate with a team at the forefront of AI research

*Certain perks and benefits are limited to full‑time employees only

You.com participates in E-Verify. We will provide the Social Security Administration (SSA) and, if necessary, the Department of Homeland Security (DHS) with information from each new employee’s Form I-9 to confirm work authorization. (English/Spanish: E-Verify Participation / Right to Work )We are also an inclusive, equitable, and accessible workplace. Please let us know if you require accommodation for any portion of the recruitment and hiring process.

Beware of Recruiting Scams:

You.com will only contact you through official @ You.com email addresses and will never ask for payment or sensitive personal information during the hiring process.

Use of AI Tools During Our Hiring Process:

At You.com, we’re thoughtful about how we use AI throughout our business — including our hiring process.

During certain stages of the interview process, we may use AI-powered tools to assist with administrative tasks such as scheduling, note‑taking, transcription, or summarizing interview discussions. These tools are intended to support our interviewers by improving accuracy and efficiency — they do not make hiring decisions or replace human judgment.

All employment decisions are made by our recruiting team and hiring managers based on a holistic review of each candidate and applicable evaluation criteria.

If AI‑assisted note‑taking or transcription may be used during your interview, we will notify you in advance. If you prefer not to participate in an interview that uses these tools, you may opt out by informing your recruiter before your scheduled interview. We will work with you to provide an alternative interview experience where reasonably practicable.

Information collected during the hiring process is handled in accordance with our Privacy Policy and applicable law. By continuing with the application process, you acknowledge that you have read this notice. If AI‑assisted tools are used during your interview and you have not opted out after receiving notice, you consent to their use as described above, where permitted by applicable law.

Voluntary Self‑Identification

For government reporting purposes, we ask candidates to respond to the below self‑identification survey.Completion of the form is entirely voluntary. Whatever your decision, it will not be considered in the hiringprocess or thereafter. Any information that you do provide will be recorded and maintained in aconfidential file.

As set forth in You.com’s Equal Employment Opportunity policy,we do not discriminate on the basis of any protected group status under any applicable law.

If you believe you belong to any of the categories of protected veterans listed below, please indicate by making the appropriate selection.As a government contractor subject to the Vietnam Era Veterans Readjustment Assistance Act (VEVRAA), we request this information in order to measurethe effectiveness of the outreach and positive recruitment efforts we undertake pursuant to VEVRAA. Classification of protected categoriesis as follows:

A "disabled veteran" is one of the following: a veteran of the U.S. military, ground, naval or air service who is entitled to compensation (or who but for the receipt of military retired pay would be entitled to compensation) under laws administered by the Secretary of Veterans Affairs; or a person who was discharged or released from active duty because of a service‑connected disability.

A "recently separated veteran" means any veteran during the three‑year period beginning on the date of such veteran's discharge or release from active duty in the U.S. military, ground, naval, or air service.

An "active duty wartime or campaign badge veteran" means a veteran who served on active duty in the U.S. military, ground, naval or air service during a war, or in a campaign or expedition for which a campaign badge has been authorized under the laws administered by the Department of Defense.

An "Armed forces service medal veteran" means a veteran who, while serving on active duty in the U.S. military, ground, naval or air service, participated in a United States military operation for which an Armed Forces service medal was awarded pursuant to Executive Order 12985.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Data Engineer
Data Engineer

Mixpeek • San Francisco (CA)

On-site
USD 180,000 - 220,000
Hubs in SF/NYC
Flexible PTO
Health insurance
+6
Data Engineer
Data Engineer

You.com • San Francisco (CA)

On-site
USD 180,000 - 220,000
Hubs in SF & NYC
Flexible PTO + US holidays
Comprehensive health insurance
+3
Data Scientist, Next Gen Recommendation Systems New York City
Data Scientist, Next Gen Recommendation Systems New York City

Impact • Northern (KY), New York (NY)

Hybrid
USD 100,000 - 125,000
Medical insurance
Dental insurance
Vision insurance
+6
Senior Developer Relations Engineer
Senior Developer Relations Engineer

You.com • San Francisco (CA)

Hybrid
USD 230,000 - 260,000
Hubs in San Francisco and New YorkCity
Flexible PTO with U.S. holidays-observ
Health insurance plan covers 100% for
+6
Senior Developer Relations Engineer
Senior Developer Relations Engineer

Mixpeek • San Francisco (CA)

On-site
USD 230,000 - 260,000
Offices SF/NYC
Flexible PTO
Health insurance
+6
Developer Community Manager
Developer Community Manager

Mixpeek • San Francisco (CA), Northern (KY)

Hybrid
USD 180,000 - 220,000
Hubs in SF & NYC
Flexible PTO
Health insurance
+5
Senior AI Engineer, Product Engineering
Senior AI Engineer, Product Engineering

Alumni Ventures • San Francisco (CA)

On-site
USD 200,000 - 250,000
Flexible PTO
Health insurance covering 100% of policyholder
401k program with 3% match
+3
Associate QA Engineer Santa Barbara, CA
Associate QA Engineer Santa Barbara, CA

Impact • Santa Barbara (CA), Northern (KY)

Hybrid
USD 70,000 - 82,000
Medical insurance
Dental insurance
Vision insurance
+11
Staff Software Engineer, AI
Staff Software Engineer, AI

Yipit, Inc. • New York (NY)

On-site
USD 180,000 - 200,000
Flexible work hours
Remote-friendly environment
401K match
+3
Machine Learning Engineer, Search Quality Mountain View, CA
Machine Learning Engineer, Search Quality Mountain View, CA

Startups • Mountain View (CA)

Hybrid
USD 159,000 - 265,000
Home office stipend
Education stipend
Wellness stipend
+1