Staff Data Platform Engineer - Petabyte-Scale Pipelines

Kindredventures

San Francisco (CA)

On-site

USD 130,000 - 170,000

Full time

14 days+
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Job summary

Kindredventures is building a Large Physics foundation Model and seeks a data engineer to design and operate the data platform for petabyte-scale storage, real-time and batch queries, and scalable ingestion.

You will own the compute and orchestration platforms (Spark, Ray, Beam), optimize data loading to training, and implement cataloging, lineage, and monitoring across the data lifecycle. Collaboration across research and engineering is essential.

Qualifications

  • Demonstrated experience building large-scale data pipelines and distributed compute systems (e.g. Spark, Ray, Beam)
  • Knowledge of state-of-the-art data ingestion, storage, and loading — including file formats and storage systems (e.g. Parquet, Zarr, Delta Lake) and how they impact performance and scalability
  • Deep familiarity with cloud infrastructure, data lake architectures, and batch and streaming pipelines
  • Understanding of how data loading throughput affects large-scale training, and experience optimizing it
  • Owns deliverables end-to-end, from collecting and translating requirements to autonomously driving execution

Responsibilities

  • Design and operate petabyte-scale storage: lakehouse architecture, file formats, and data layout optimized for both batch and real-time queries
  • Own the shared compute and orchestration platform (e.g. Spark, Ray, workflow scheduling) that ingestion and research pipelines run on
  • Optimize data strategy end to end from storage to loading, owning high-throughput data loading into training up to the tensor boundary
  • Build systems for cataloging, deduplication, lineage, search, and reproducibility at every stage of the data lifecycle
  • Implement the platform-level quality and monitoring tooling that data and research teams build their checks on
  • Scale infrastructure to improve engineering velocity and ensure reliability, with monitoring and alerting to match
  • Work across the full data lifecycle when the mission needs it — including building and operating ingestion pipelines for critical data sources directly

Skills

Large-scale data pipelines
Distributed compute systems
Cloud infrastructure
Data ingestion best practices
Autonomous delivery ownership

Tools

Spark
Ray
Beam
Parquet
Zarr
Delta Lake

Job description

Kindredventures is building a Large Physics foundation Model and seeks a data engineer to design and operate the data platform for petabyte-scale storage, real-time and batch queries, and scalable ingestion.

You will own the compute and orchestration platforms (Spark, Ray, Beam), optimize data loading to training, and implement cataloging, lineage, and monitoring across the data lifecycle. Collaboration across research and engineering is essential.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Staff Data Infrastructure Engineer (Petabyte-scale)
Staff Data Infrastructure Engineer (Petabyte-scale)

Causal Labs • San Francisco (CA)

On-site
USD 150,000 - 210,000
Data Ingestion & Quality Engineer — End-to-End Pipelines
Data Ingestion & Quality Engineer — End-to-End Pipelines

Kindredventures • San Francisco (CA)

On-site
USD 120,000 - 180,000
Senior Data Platform Engineer: Scale Petabyte Pipelines
Senior Data Platform Engineer: Scale Petabyte Pipelines

Weights & Biases • Livingston (NJ)

On-site
USD 165,000 - 242,000
Senior Data Infrastructure Engineer, Petabyte Pipelines
Senior Data Infrastructure Engineer, Petabyte Pipelines

Genesis AI • United States

On-site
USD 140,000 - 210,000
Member of Technical Staff — Data Infrastructure
Member of Technical Staff — Data Infrastructure

Causal Labs • San Francisco (CA)

On-site
USD 150,000 - 210,000
Staff Data Platform Engineer: Pipelines & AI
Staff Data Platform Engineer: Pipelines & AI

Perplexity • Palo Alto (CA)

On-site
USD 220,000 - 405,000
Member of Technical Staff — Data Infrastructure
Member of Technical Staff — Data Infrastructure

Kindredventures • San Francisco (CA)

On-site
USD 130,000 - 170,000
Data Platform Engineering Lead — Build & Scale Pipelines
Data Platform Engineering Lead — Build & Scale Pipelines

Sierra Ventures • San Francisco (CA)

On-site
USD 260,000 - 380,000
Stock options
Health & wellness benefits
Meals provided in the office
+1
Senior Data Platform Engineer — AI-Scale Pipelines
Senior Data Platform Engineer — AI-Scale Pipelines

Weights & Biases • New York (NY)

On-site
USD 165,000 - 242,000
Medical, dental, vision insurance
401(k) with employer match
Flexible PTO
+2
Staff Data Platform Engineer — Scalable Data Foundations
Staff Data Platform Engineer — Scalable Data Foundations

Parallel • San Francisco (CA)

On-site
USD 180,000 - 240,000
Competitive salary
Generous equity
Visa sponsorships
+5