Staff Data Infra Architect – High-Scale ML Pipelines

Rhoda AI

Palo Alto (CA)

On-site

USD 180,000 - 280,000

Full time

14 days+
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Job summary

Rhoda AI in Palo Alto seeks a senior data infrastructure leader to architect, build, and scale a high-throughput data platform handling billions of video clips. You will own systems from storage to indexing, focusing on reliability, latency, and cost efficiency while collaborating with research and ML teams.

You will design observable pipelines, optimize throughput, manage data artifacts and lineage, and develop internal tools to enable researchers to explore datasets at scale.

Qualifications

  • 5+ years in data infrastructure, distributed systems, ML infrastructure, or related field.
  • Experience building and operating large-scale data pipelines (1B+ samples or petabyte-scale).
  • Strong understanding of distributed databases, storage, and cloud architectures.
  • Proven ability to optimize throughput, balance workloads, and manage costs in cloud environments.
  • Experience with Ray or Spark for large-scale processing and transformation.
  • Strong observability, monitoring, and reliability practices for high-scale systems.
  • Ability to own end-to-end systems from design to production.
  • Staff-level candidates should define technical direction and own architectural decisions.

Responsibilities

  • Architect, build, and scale high-throughput data infrastructure for billions of video clips with reliability, latency, and cost guarantees.
  • Design and optimize large-scale storage systems for multimodal datasets.
  • Build efficient indexing and retrieval systems for fast dataset querying and iteration.
  • Develop observability frameworks for data pipelines including monitoring and failure recovery.
  • Implement intelligent workload balancing across distributed compute and storage.
  • Manage data artifacts, versioning, and lineage for reproducibility across training runs.
  • Build internal tools to enable researchers and engineers to explore and analyze large datasets at scale.
  • Support integration of vision-language models within data pipelines for screening and metadata generation.

Skills

Data infrastructure
Distributed systems
ML infrastructure
Cloud storage architectures
Observability & reliability
End-to-end ownership

Education

Bachelor's degree in Computer Science or related

Tools

Ray
Spark
BigQuery
Redshift

Job description

Rhoda AI in Palo Alto seeks a senior data infrastructure leader to architect, build, and scale a high-throughput data platform handling billions of video clips. You will own systems from storage to indexing, focusing on reliability, latency, and cost efficiency while collaborating with research and ML teams.

You will design observable pipelines, optimize throughput, manage data artifacts and lineage, and develop internal tools to enable researchers to explore datasets at scale.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Staff Data Infrastructure Architect
Staff Data Infrastructure Architect

Rhoda AI • Mountain View (CA)

On-site
USD 180,000 - 240,000
Staff Research Engineer, Scalable Video Data & Evaluation
Staff Research Engineer, Scalable Video Data & Evaluation

Rhoda AI • Mountain View (CA)

On-site
USD 120,000 - 160,000
Senior Data Engineer: Scalable Pipelines & Data Platforms
Senior Data Engineer: Scalable Pipelines & Data Platforms

Rad AI • San Francisco (CA)

On-site
USD 145,000 - 190,000
Medical, Dental, Vision & Life
HSA with employer match
401(k)
+4
Research Member of Technical Staff- Data Infrastructure
Research Member of Technical Staff- Data Infrastructure

Rhoda AI • Mountain View (CA)

On-site
USD 180,000 - 240,000
Senior Data Infra Architect - Scale ML Pipelines
Senior Data Infra Architect - Scale ML Pipelines

Archer56 • San Jose (CA)

On-site
USD 130,000 - 160,000
Data Engineering Lead: Scalable ML Data Pipelines
Data Engineering Lead: Scalable ML Data Pipelines

Hark • San Jose (CA)

On-site
USD 170,000 - 450,000
Senior Data Engineer - Healthcare AI Pipelines
Senior Data Engineer - Healthcare AI Pipelines

Engg • San Francisco (CA)

Hybrid
USD 150,000 - 210,000
Medical Insurance
Dental Insurance
Vision Insurance
+3
Senior Data Infrastructure Engineer - Real-Time Pipelines
Senior Data Infrastructure Engineer - Real-Time Pipelines

Decagon • San Francisco (CA)

On-site
USD 200,000 - 400,000
Medical, Dental, and Vision benefits
Retirement Plan (401K)
Parental Leave
+4
Senior Data Platform Engineer for Scalable AI Pipelines
Senior Data Platform Engineer for Scalable AI Pipelines

Ouster • San Francisco (CA)

On-site
USD 180,000 - 240,000
Data Foundations Engineer – Scalable Pipelines & Open AI
Data Foundations Engineer – Scalable Pipelines & Open AI

Reflection • San Francisco (CA)

On-site
USD 150,000 - 210,000
Top-tier compensation
Stock options
Health & wellness
+3