Member of Technical Staff (Data Intelligence)

Reka

United States

On-site

USD 100,000 - 130,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Reka is seeking a Data Engineer to ensure high-quality data production at petabyte scale. You will collaborate with model researchers and build algorithms for automated data quality assessment while tracking datasets for reproducibility. The ideal candidate has strong machine learning and deep learning fundamentals, practical experience with distributed processing tools, and solid Python skills. This position requires a blend of research and production engineering, offering a dynamic work environment in the United States.

Qualifications

  • Strong fundamentals in ML and experience with large-scale systems.
  • Comfortable with both research and production engineering.
  • Demonstrated experience with data quality and dataset releases.
  • Ability to design experiments with unbiased outcomes.
  • Practical experience with distributed processing tools.

Responsibilities

  • Define quality metrics, validation checks, and acceptance thresholds for data.
  • Create internal datasets for building fundamental World Models.
  • Build algorithms for automated data quality assessment.
  • Track datasets, metadata, and ensure experiments are reproducible.
  • Own CI/CD for the data stack and automate workflows.

Skills

Machine Learning fundamentals
Deep learning experience
Data quality assessment
Python skills
Distributed processing
GitHub experience

Tools

PyTorch
Spark
Airflow

Job description

In this role, you’ll work closely with model researchers, data infrastructure engineers, and cross-functional partners to make sure our data is high quality and can be produced at petabyte scale in a reliable, efficient way. From understanding how data choices show up in model behavior, to building processing pipelines and running the compute behind them, you’ll help ensure our models are trained on the best data we can get.

What you’ll do
  • Work with model researchers to define what “good data” means for our models, including quality metrics, validation checks, and acceptance thresholds

  • Explore open source datasets and create internal ones most suitable to build fundamental World Models

  • Build algorithms for automated data quality assessment, data domain mixtures, and domain adaptation from synthetic to real data.

  • Track datasets, metadata, provenance, and versions so experiments are reproducible and it’s clear what data went into which training and evaluation runs

  • Own CI/CD and development tooling for the data stack (GitHub, Python, PyTorch), and automate repetitive workflows to reduce friction

  • Track and optimize throughput, storage, and compute utilization across pipelines and related assets

What we’re looking for
  • Strong ML and deep learning fundamentals with experience building and operating large-scale data and/or compute systems

  • Comfortable moving between research questions and production engineering: you can dig into data, run analyses, and also ship reliable systems

  • Demonstrated research experience with data compositions, quality, and dataset releases

  • Ability to design and execute experiments with convincing unbiased outcomes

  • Practical experience with distributed processing and orchestration (Spark, Ray, Airflow, or equivalents)

  • Solid Python skills, and familiarity with the tooling around modern model training workflows (datasets, checkpoints, experiment tracking)

  • Strong instincts around data quality: how to measure it, how to monitor it, and how to prevent regressions as things scale

  • Able to work in a fast-moving environment, prioritize what matters, and communicate clearly with both researchers and engineers

  • Bonus: experience with large video datasets, dataset curation for training, or building internal tooling for evaluation/analysis in ML environments

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Member of Technical Staff, Data
Member of Technical Staff, Data

Inception • San Francisco (CA)

On-site
USD 180,000 - 240,000
Member of Technical Staff, Data Infrastructure
Member of Technical Staff, Data Infrastructure

Inception • San Francisco (CA)

On-site
USD 140,000 - 190,000
Member of Technical Staff — Data Ingestion & Quality
Member of Technical Staff — Data Ingestion & Quality

Kindredventures • San Francisco (CA)

On-site
USD 120,000 - 180,000
Member of Technical Staff — Data Ingestion & Quality
Member of Technical Staff — Data Ingestion & Quality

Causal • San Francisco (CA)

On-site
USD 120,000 - 160,000
Research Member of Technical Staff- Data Infrastructure
Research Member of Technical Staff- Data Infrastructure

Rhoda AI • Palo Alto (CA)

On-site
USD 180,000 - 280,000
Member of Technical Staff — Data Ingestion & Quality
Member of Technical Staff — Data Ingestion & Quality

Causal Labs • San Francisco (CA)

On-site
USD 120,000 - 170,000
Member of Technical Staff — Data Infrastructure
Member of Technical Staff — Data Infrastructure

Kindredventures • San Francisco (CA)

On-site
USD 130,000 - 170,000
Member of Technical Staff — Data Infrastructure
Member of Technical Staff — Data Infrastructure

Causal • San Francisco (CA)

On-site
USD 150,000 - 190,000
Member of Technical Staff — Data Infrastructure
Member of Technical Staff — Data Infrastructure

Causal Labs • San Francisco (CA)

On-site
USD 150,000 - 210,000
Member of Technical Staff - ML Research
Member of Technical Staff - ML Research

Kindredventures • San Francisco (CA)

On-site
USD 120,000 - 160,000