Data Intelligence Engineer, Scale-Out ML Data Pipelines

reka

Singapore

Remote

SGD 90,000 - 150,000

Full time

14 days+
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Job summary

Reka is a globally distributed foundation model startup, headquartered in the San Francisco Bay Area, California. Embracing a remote-first approach, our team brings together top talent from around the world.

Our founding team, along with many of our team members, has contributed to many of the breakthroughs in AI over the past decade. In this role, you’ll work with model researchers, data infrastructure engineers, and cross-functional partners to ensure data is high quality and can be produced

Qualifications

  • Strong ML and deep learning fundamentals with experience building and operating large-scale data and/or compute systems.
  • Comfortable moving between research questions and production engineering: you can dig into data, run analyses, and also ship reliable systems
  • Demonstrated research experience with data compositions, quality, and dataset releases
  • Ability to design and execute experiments with convincing unbiased outcomes
  • Practical experience with distributed processing and orchestration (Spark, Ray, Airflow, or equivalents)
  • Solid Python skills, and familiarity with the tooling around modern model training workflows (datasets, checkpoints, experiment tracking)
  • Strong instincts around data quality: how to measure it, how to monitor it, and how to prevent regressions as things scale
  • Able to work in a fast-moving environment, prioritize what matters, and communicate clearly with both researchers and engineers
  • Bonus: experience with large video datasets, dataset curation for training, or building internal tooling for evaluation/analysis in ML environments

Responsibilities

  • Define what “good data” means for our models, including quality metrics, validation checks, and acceptance thresholds
  • Explore open source datasets and create internal ones suited for World Models
  • Build algorithms for automated data quality assessment, data domain mixtures, and domain adaptation from synthetic to real data
  • Track datasets, metadata, provenance, and versions for reproducible experiments
  • Own CI/CD and development tooling for the data stack (GitHub, Python, PyTorch), and automate repetitive workflows to reduce friction
  • Track and optimize throughput, storage, and compute utilization across pipelines and related assets

Skills

ML fundamentals
Python skills
Distributed processing
Experiment design
Data quality
Production engineering
Spark
Airflow
Ray

Tools

Spark
Ray
Airflow
GitHub
PyTorch

Job description

Reka is a globally distributed foundation model startup, headquartered in the San Francisco Bay Area, California. Embracing a remote-first approach, our team brings together top talent from around the world.

Our founding team, along with many of our team members, has contributed to many of the breakthroughs in AI over the past decade. In this role, you’ll work with model researchers, data infrastructure engineers, and cross-functional partners to ensure data is high quality and can be produced

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Member of Technical Staff (Data Intelligence)
Member of Technical Staff (Data Intelligence)

reka • Singapore

Remote
SGD 90,000 - 150,000
Global Data Platform Engineer — Real-Time Pipelines & ML
Global Data Platform Engineer — Real-Time Pipelines & ML

re-zoo-me • Singapore

Hybrid
SGD 120,000 - 180,000
Senior Data Engineering Consultant: Real-Time Pipelines & AI
Senior Data Engineering Consultant: Real-Time Pipelines & AI

re-zoo-me • Singapore

On-site
SGD 120,000 - 190,000
Senior Data Engineer:Scalable Pipelines, Lakehouse & DevOps
Senior Data Engineer:Scalable Pipelines, Lakehouse & DevOps

re-zoo-me • Singapore

Hybrid
SGD 180,000 - 240,000
Senior Data Engineer - Scalable Platform & Warehouse
Senior Data Engineer - Scalable Platform & Warehouse

re-zoo-me • Singapore

Hybrid
SGD 150,000 - 210,000
Senior Data Engineer — Remote, Build Scalable Analytics & ML
Senior Data Engineer — Remote, Build Scalable Analytics & ML

Kake • Singapore

On-site
SGD 153,000 - 230,000
Competitive USD pay
Fully remote work
Better Me Fund
+1
Senior Data & Analytics Engineer - Pipelines & ETL
Senior Data & Analytics Engineer - Pipelines & ETL

re-zoo-me • Singapore

Hybrid
SGD 90,000 - 130,000
Senior Data & Analytics Engineer - Enterprise Pipelines
Senior Data & Analytics Engineer - Enterprise Pipelines

re-zoo-me • Singapore

Hybrid
SGD 90,000 - 150,000
Senior Data Engineer: AWS Data Platform & Pipelines
Senior Data Engineer: AWS Data Platform & Pipelines

re-zoo-me • Singapore

Hybrid
SGD 120,000 - 180,000
Senior Data Platform Engineer - ML Data Pipelines
Senior Data Platform Engineer - ML Data Pipelines

WD • Singapore

On-site
SGD 90,000 - 140,000