Data Infra Engineer: Scale Distributed Data Pipelines

Mosaic.tech

San Francisco (CA)

On-site

USD 350,000 - 475,000

Full time

14 days+
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Benefits offered by this job

Health, dental, and vision benefits
Unlimited PTO
Paid parental leave
Relocation support

Job summary

Thinking Machines is seeking an engineer to help design and scale the core data infrastructure behind distributed training pipelines and multimodal data catalogs. You’ll work with researchers to accelerate experiments, improve data quality, and build systems that scale across petabytes of data.

Ideal candidates have hands-on Python or Rust experience and comfort with Spark or Ray, cloud data lakes, and end-to-end project ownership, in a collaborative, fast-moving environment in San Francisco.

Qualifications

  • Bachelor’s degree or equivalent in computer science, engineering, or similar.
  • Proficiency in Python or Rust.
  • Fluent in distributed compute frameworks such as Apache Spark or Ray.
  • Familiar with cloud infrastructure, data lake architectures, and batch/streaming pipelines.
  • Comfort operating across the stack and owning projects end-to-end.
  • Thrive in a highly collaborative environment with cross-functional partners.
  • Proactive with initiative to work across stacks to ship.

Responsibilities

  • Design, build, and operate scalable, fault-tolerant infrastructure for LLM Research: distributed compute, data orchestration, and storage across modalities.
  • Develop high-throughput systems for data ingestion, processing, and transformation — including training data catalogs, deduplication, quality checks, and search.
  • Build systems for traceability, reproducibility, and robust quality control at every stage of the data lifecycle.
  • Implement and maintain monitoring and alerting to support platform reliability and performance.
  • Collaborate with research teams to unlock new features, improve data quality, and accelerate training cycles.

Job description

Thinking Machines is seeking an engineer to help design and scale the core data infrastructure behind distributed training pipelines and multimodal data catalogs. You’ll work with researchers to accelerate experiments, improve data quality, and build systems that scale across petabytes of data.

Ideal candidates have hands-on Python or Rust experience and comfort with Spark or Ray, cloud data lakes, and end-to-end project ownership, in a collaborative, fast-moving environment in San Francisco.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Staff Data Infrastructure Engineer: Scale AI Pipelines
Staff Data Infrastructure Engineer: Scale AI Pipelines

Inception • San Francisco (CA)

On-site
USD 140,000 - 190,000
Software Engineer, Data Infra
Software Engineer, Data Infra

Mosaic.tech • San Francisco (CA)

On-site
USD 350,000 - 475,000
Health, dental, and vision benefits
Unlimited PTO
Paid parental leave
+1
Data Infrastructure Engineer — Scalable ML Data Pipelines | Flexible Hours
Data Infrastructure Engineer — Scalable ML Data Pipelines | Flexible Hours

JobCubby • San Francisco (CA), Northern (KY)

Hybrid
USD 500,000 - 850,000
Visa sponsorship
Office in San Francisco
Flexible working hours
Remote Data Engineer — Scalable Pipelines & AI Data
Remote Data Engineer — Scalable Pipelines & AI Data

Crossing Hurdles • United States

On-site
USD 140,000 - 180,000
Remote AI Data Engineer: Scale Data Pipelines for ML
Remote AI Data Engineer: Scale Data Pipelines for ML

Bright Vision Technologies • Farmington Hills (MI)

On-site
USD 80,000 - 100,000
Data Infrastructure Engineer — Scale & Reliability
Data Infrastructure Engineer — Scale & Reliability

OpenAI • California (MO)

Hybrid
USD 150,000 - 190,000
Relocation assistance
Data Movement Platform Engineer - Scale Data Pipelines
Data Movement Platform Engineer - Scale Data Pipelines

Reddit, Inc. • Los Angeles (CA)

On-site
USD 164,000 - 230,000
Comprehensive Healthcare Benefits
401k with Employer Match
Global Benefits Programs
+3
Data Foundations Engineer – Scalable Pipelines & Open AI
Data Foundations Engineer – Scalable Pipelines & Open AI

Reflection • San Francisco (CA)

On-site
USD 150,000 - 210,000
Top-tier compensation
Stock options
Health & wellness
+3
Remote AI Data Engineer - Scale-Pipeline Expert
Remote AI Data Engineer - Scale-Pipeline Expert

Bright-Vision-Technologies • United States

Remote
USD 80,000 - 100,000
Member of Technical Staff, Data Infrastructure
Member of Technical Staff, Data Infrastructure

Inception • San Francisco (CA)

On-site
USD 140,000 - 190,000