Staff Data Engineer - Synthetic Data & AI Pipelines

Inception

San Francisco (CA)

On-site

USD 180,000 - 240,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Inception in San Francisco is seeking experienced engineers and scientists to shape how we collect, process, and curate datasets powering our models. You will build scalable data pipelines, develop synthetic data generation techniques, and ensure models train on high-quality, diverse data.

The role combines engineering with research, requiring hands-on implementation with Python, Spark, and ML frameworks. You will design data ingestion systems, manage large-scale storage, and create evaluation

Qualifications

  • 3+ years of experience building data processing pipelines at scale for AI/ML.
  • Strong proficiency in Python and experience with data processing frameworks (Spark, Beam, Airflow).
  • Familiarity with synthetic data generation techniques and data augmentation strategies.
  • Familiarity with web scraping, crawling technologies, and Common Crawl datasets.
  • Solid understanding of machine learning fundamentals and experience with ML frameworks (PyTorch, TensorFlow).
  • Experience with SQL and NoSQL databases for managing structured and unstructured data.

Responsibilities

  • Develop data mixes for training LLMs, including leveraging open-source datasets, synthetically generated data, and curated human feedback.
  • Design and implement data pipelines for processing petabyte-scale datasets.
  • Build systems for web crawling, data ingestion, and real-time data processing to support model training.
  • Develop tools and frameworks for efficient data storage, retrieval, and versioning across distributed systems.
  • Create evaluation frameworks to measure data diversity, quality, and representativeness.
  • Ensure data collection adheres to privacy regulations.

Skills

Python
Data pipelines
ML fundamentals

Education

BS/MS/PhD in CS/ML or related field

Tools

Apache Spark
Beam
Airflow
PyTorch
TensorFlow
SQL
NoSQL
Common Crawl
S3
BigQuery

Job description

Inception in San Francisco is seeking experienced engineers and scientists to shape how we collect, process, and curate datasets powering our models. You will build scalable data pipelines, develop synthetic data generation techniques, and ensure models train on high-quality, diverse data.

The role combines engineering with research, requiring hands-on implementation with Python, Spark, and ML frameworks. You will design data ingestion systems, manage large-scale storage, and create evaluation

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Staff Data Infrastructure Engineer: Scale AI Pipelines
Staff Data Infrastructure Engineer: Scale AI Pipelines

Inception • San Francisco (CA)

On-site
USD 140,000 - 190,000
Synthetic Data Engineer for AI Training Pipelines
Synthetic Data Engineer for AI Training Pipelines

Hyphen Connect Limited • Seattle (WA)

On-site
USD 100,000 - 130,000
Senior ML Engineer: Synthetic Data Pipelines
Senior ML Engineer: Synthetic Data Pipelines

Cohere • United States

Remote
USD 180,000 - 260,000
Production-Grade AI & Synthetic Data Engineer
Production-Grade AI & Synthetic Data Engineer

TypeSafe AI • San Francisco (CA)

On-site
USD 180,000 - 280,000
Base salary of $180k-$280k plus equity
100% covered health insurance
Daily lunch and dinner
+2
Senior Data Engineer: Build Scalable Data Pipelines
Senior Data Engineer: Build Scalable Data Pipelines

Pattern AI • Hayward Park (CA)

On-site
USD 100,000 - 130,000
Synthetic Data Engineer (AI Data/Training)
Synthetic Data Engineer (AI Data/Training)

Hyphen Connect Limited • Seattle (WA)

On-site
USD 100,000 - 130,000
Senior Data Engineer: Scalable Pipelines & Data Platforms
Senior Data Engineer: Scalable Pipelines & Data Platforms

Rad AI • San Francisco (CA)

On-site
USD 145,000 - 190,000
Medical, Dental, Vision & Life
HSA with employer match
401(k)
+4
Staff Data Engineering Lead - Scalable Pipelines
Staff Data Engineering Lead - Scalable Pipelines

EngineersOfAI • Seattle (WA)

On-site
USD 130,000 - 170,000
Data Foundations Engineer – Scalable Pipelines & Open AI
Data Foundations Engineer – Scalable Pipelines & Open AI

Reflection • San Francisco (CA)

On-site
USD 150,000 - 210,000
Top-tier compensation
Stock options
Health & wellness
+3
Staff Data Pipelines Engineer for AI-Driven Finance
Staff Data Pipelines Engineer for AI-Driven Finance

Autonomous Technologies Group • San Francisco (CA)

On-site
USD 140,000 - 210,000