Staff Data Infrastructure Engineer for Large-Scale ML

Rhoda AI

Mountain View (WY)

On-site

USD 180,000 - 260,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Rhoda AI is seeking Data Infrastructure MLEs to scale our data pipelines and model-training data infrastructure for billions of video clips. You will design high-throughput storage, indexing, and retrieval systems for multimodal datasets, while ensuring reliability, observability, and cost efficiency.

You will own end-to-end data systems from ingestion to production, implement workload balancing across distributed compute and storage, and collaborate with researchers to enable scalable data

Qualifications

  • 5+ years of experience in data infrastructure, distributed systems, ML infrastructure, or related field.
  • Experience building and operating large-scale data pipelines (1B+ samples or petabyte-scale systems).
  • Deep understanding of distributed systems, databases, indexing strategies, and cloud storage architectures.
  • Experience with throughput optimization and cost-performance tradeoffs in cloud environments.
  • Familiarity with Ray or Spark for large-scale data processing.
  • Strong observability, monitoring, and production reliability capabilities.
  • Ability to own systems end-to-end from design to production.

Responsibilities

  • Architect, build, and scale a high-throughput data infrastructure that processes and manages billions of video clips with strong guarantees around reliability, latency, and cost efficiency
  • Design and optimize large-scale storage systems (cloud object storage, databases, metadata stores) for multimodal datasets
  • Build efficient indexing and retrieval systems to support fast dataset querying, filtering, and iteration for research and production use cases
  • Develop observability frameworks for data pipelines including monitoring, alerting, failure recovery, and performance optimization
  • Implement intelligent workload balancing and throughput optimization across distributed compute and storage systems
  • Manage data artifacts, versioning, and lineage to ensure reproducibility and traceability across training runs
  • Build internal interfaces and lightweight tools that enable researchers and engineers to explore, query, and analyze large datasets at scale
  • Support integration and scalable deployment of vision-language models (VLMs) within data pipelines for screening, enrichment, or metadata generation

Skills

Data infrastructure
Distributed systems
ML infrastructure
Large-scale data pipelines
Cloud storage architectures
Observability & reliability
End-to-end ownership

Tools

Ray
Spark
Snowflake
BigQuery
Kubernetes

Job description

Rhoda AI is seeking Data Infrastructure MLEs to scale our data pipelines and model-training data infrastructure for billions of video clips. You will design high-throughput storage, indexing, and retrieval systems for multimodal datasets, while ensuring reliability, observability, and cost efficiency.

You will own end-to-end data systems from ingestion to production, implement workload balancing across distributed compute and storage, and collaborate with researchers to enable scalable data

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Staff Data Infra Architect – High-Scale ML Pipelines
Staff Data Infra Architect – High-Scale ML Pipelines

Rhoda AI • Palo Alto (CA)

On-site
USD 180,000 - 280,000
Staff Data Infrastructure Architect
Staff Data Infrastructure Architect

Rhoda AI • Mountain View (CA)

On-site
USD 180,000 - 240,000
Research Member of Technical Staff- Data Infrastructure
Research Member of Technical Staff- Data Infrastructure

Rhoda AI • Mountain View (WY)

On-site
USD 180,000 - 260,000
Staff Research Engineer, Scalable Video Data & Evaluation
Staff Research Engineer, Scalable Video Data & Evaluation

Rhoda AI • Mountain View (CA)

On-site
USD 120,000 - 160,000
Lead ML Training Systems Engineer - Multimodal, Large-Scale
Lead ML Training Systems Engineer - Multimodal, Large-Scale

Rhoda AI • Palo Alto (CA)

On-site
USD 210,000 - 320,000
Senior ML Systems Engineer - Robot Learning Pipeline
Senior ML Systems Engineer - Robot Learning Pipeline

Socket.dev • Mountain View (CA)

On-site
USD 180,000 - 320,000
Research Member of Technical Staff- Data Infrastructure
Research Member of Technical Staff- Data Infrastructure

Rhoda AI • Mountain View (CA)

On-site
USD 180,000 - 240,000
ML Data Infrastructure Engineer
ML Data Infrastructure Engineer

Physical Intelligence • San Francisco (CA)

On-site
USD 180,000 - 240,000
Senior ML Training Systems Engineer, Large-Scale Robotics
Senior ML Training Systems Engineer, Large-Scale Robotics

Rhoda AI • Mountain View (CA)

On-site
USD 150,000 - 200,000
Member of Technical Staff, Data Infrastructure
Member of Technical Staff, Data Infrastructure

Inception • San Francisco (CA)

On-site
USD 140,000 - 190,000