Research Member of Technical Staff- Data Infrastructure

Rhoda AI

Palo Alto (CA)

On-site

USD 180,000 - 280,000

Full time

14 days+
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Job summary

Rhoda AI in Palo Alto seeks a senior data infrastructure leader to architect, build, and scale a high-throughput data platform handling billions of video clips. You will own systems from storage to indexing, focusing on reliability, latency, and cost efficiency while collaborating with research and ML teams.

You will design observable pipelines, optimize throughput, manage data artifacts and lineage, and develop internal tools to enable researchers to explore datasets at scale.

Qualifications

  • 5+ years in data infrastructure, distributed systems, ML infrastructure, or related field.
  • Experience building and operating large-scale data pipelines (1B+ samples or petabyte-scale).
  • Strong understanding of distributed databases, storage, and cloud architectures.
  • Proven ability to optimize throughput, balance workloads, and manage costs in cloud environments.
  • Experience with Ray or Spark for large-scale processing and transformation.
  • Strong observability, monitoring, and reliability practices for high-scale systems.
  • Ability to own end-to-end systems from design to production.
  • Staff-level candidates should define technical direction and own architectural decisions.

Responsibilities

  • Architect, build, and scale high-throughput data infrastructure for billions of video clips with reliability, latency, and cost guarantees.
  • Design and optimize large-scale storage systems for multimodal datasets.
  • Build efficient indexing and retrieval systems for fast dataset querying and iteration.
  • Develop observability frameworks for data pipelines including monitoring and failure recovery.
  • Implement intelligent workload balancing across distributed compute and storage.
  • Manage data artifacts, versioning, and lineage for reproducibility across training runs.
  • Build internal tools to enable researchers and engineers to explore and analyze large datasets at scale.
  • Support integration of vision-language models within data pipelines for screening and metadata generation.

Skills

Data infrastructure
Distributed systems
ML infrastructure
Cloud storage architectures
Observability & reliability
End-to-end ownership

Education

Bachelor's degree in Computer Science or related

Tools

Ray
Spark
BigQuery
Redshift

Job description

What\'s You'll Do
  • Architect, build, and scale a high-throughput data infrastructure that processes and manages billions of video clips with strong guarantees around reliability, latency, and cost efficiency

  • Design and optimize large-scale storage systems (cloud object storage, databases, metadata stores) for multimodal datasets

  • Build efficient indexing and retrieval systems to support fast dataset querying, filtering, and iteration for research and production use cases

  • Develop observability frameworks for data pipelines including monitoring, alerting, failure recovery, and performance optimization

  • Implement intelligent workload balancing and throughput optimization across distributed compute and storage systems

  • Manage data artifacts, versioning, and lineage to ensure reproducibility and traceability across training runs

  • Build internal interfaces and lightweight tools that enable researchers and engineers to explore, query, and analyze large datasets at scale

  • Support integration and scalable deployment of vision-language models (VLMs) within data pipelines for screening, enrichment, or metadata generation

What\'s We're Looking For
  • 5+ years of experience in data infrastructure, distributed systems, ML infrastructure, or a closely related field

  • Strong experience building and operating large-scale data pipelines (1B+ samples or petabyte-scale systems preferred)

  • Deep understanding of distributed systems, databases, indexing strategies, and cloud storage architectures

  • Experience optimizing data throughput, workload balancing, and cost-performance tradeoffs in cloud environments

  • Experience with distributed compute frameworks such as Ray or Spark for large-scale data processing and transformation

  • Strong skills in observability, monitoring, and production reliability for high-scale systems

  • Strong software engineering fundamentals with the ability to own systems end-to-end, from design to production

  • Staff-level candidates are expected to define technical direction and own architectural decisions independently; senior candidates execute complex systems work with strong fundamentals and growing scope

Nice to Have (But Not Required)
  • Experience managing large multimodal datasets

  • Familiarity with ML training workflows and data lifecycle management

  • Familiarity with vision-language models (VLMs) and experience running ML inference workloads at scale in distributed or cloud environments

  • Experience with robotics data formats or real-world sensor data (video, proprioception, teleoperation logs)

  • Experience with data warehouse technologies (e.g., Snowflake, BigQuery, or Redshift) for large-scale data storage, querying, and analytics

  • Familiarity with data versioning and lineage tooling (e.g., DVC, Delta Lake, or similar)

Why This Role
  • Own the data foundation that everything else runs on — model quality is only as good as the data infrastructure beneath it

  • Direct collaboration with research and ML systems teams; your work has immediate, measurable impact on training velocity

  • High ownership in a small team — you\'ll make real architectural decisions, not execute tickets

  • Help build the infrastructure that powers robots operating in the real world, at scale

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Research Member of Technical Staff- Data Infrastructure
Research Member of Technical Staff- Data Infrastructure

Rhoda AI • Mountain View (CA)

On-site
USD 180,000 - 240,000
Member of Technical Staff, Data Infrastructure
Member of Technical Staff, Data Infrastructure

Inception • San Francisco (CA)

On-site
USD 140,000 - 190,000
Member of Technical Staff — Data Infrastructure
Member of Technical Staff — Data Infrastructure

Causal Labs • San Francisco (CA)

On-site
USD 150,000 - 210,000
Member of Technical Staff, Backend
Member of Technical Staff, Backend

Nomadic AI • San Francisco (CA), Northern (KY)

On-site
USD 180,000 - 240,000
Member of Technical Staff — Data Infrastructure
Member of Technical Staff — Data Infrastructure

Kindredventures • San Francisco (CA)

On-site
USD 130,000 - 170,000
Member of Technical Staff — Data Infrastructure
Member of Technical Staff — Data Infrastructure

Causal • San Francisco (CA)

On-site
USD 150,000 - 190,000
Staff Data Engineer
Staff Data Engineer

LiveView Technologies • Seattle (WA)

On-site
USD 172,000 - 221,000
Health, dental, and vision coverage
401k match up to 4%
Flexible PTO
ML Research Engineer, Data
ML Research Engineer, Data

Weave Robotics • San Francisco (CA)

On-site
USD 140,000 - 190,000
Research Member of Technical Staff- Training Systems
Research Member of Technical Staff- Training Systems

Rhoda AI • Palo Alto (CA)

On-site
USD 210,000 - 320,000
Member of Technical Staff — Data Ingestion & Quality
Member of Technical Staff — Data Ingestion & Quality

Kindredventures • San Francisco (CA)

On-site
USD 120,000 - 180,000