Data Infrastructure

Genesis AI

Northern (KY)

Hybrid

USD 120,000 - 180,000

Full time

14 days+
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Job summary

Genesis AI is seeking a highly skilled Data Engineer to design, build, and operate data pipelines for robotics foundation model training at petabyte scale in the United States. You will own core data infrastructure, including data models, storage, ingestion, transformation, and orchestration layers.

You will standardize data models across real-world teleoperation and synthetic simulation datasets while collaborating with a motivated team to advance general-purpose Physical AI.

Qualifications

  • 8+ years of experience building large-scale data pipelines.
  • Experience with distributed systems and data storage tech.
  • Production-grade infrastructure experience.

Responsibilities

  • Design, build, and maintain data pipelines for robotics foundation model training.
  • Own data infrastructure including data model, storage, ingestion, transformation, and orchestration.
  • Standardize data models and unify processing pipelines across datasets.
  • Collaborate with team to advance general-purpose Physical AI.

Skills

Python
Go
Large-scale data pipelines
Distributed systems
Data modeling

Tools

Spark
Kafka
S3
Terraform
Kubernetes

Job description

What You’ll Do
  • Design, build, and maintain large-scale data pipelines (batch and streaming) for robotics foundation model training and evaluation at petabyte scale
  • Own core data infrastructure: data model, storage systems, ingestion pipelines, transformation frameworks, and orchestration layers
  • Standardize data models and unify processing pipelines across real-world teleoperation and synthetic simulation datasets
  • Collaborate with a team of driven individuals committed to building general-purpose Physical AI
What You’ll Bring
  • Excellent software engineering skills (Python, Go, or similar)
  • Extensive experience designing, building, and maintaining large-scale data pipelines (8+ years)
  • Deep understanding of distributed systems (Spark, Kafka, or similar)
  • Extensive experience with data storage technologies (data lakes, warehouses, object stores like S3)
  • Experience running and maintaining production-grade infrastructure (Kubernetes, Terraform)
  • Bonus: Experience supporting AI systems, in particular embodied AI like self-driving
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Data Infrastructure
Data Infrastructure

Genesis AI • United States

On-site
USD 140,000 - 210,000
Data Infrastructure Engineer
Data Infrastructure Engineer

Mind Robotics Inc. • Palo Alto (CA), Northern (KY)

Hybrid
USD 130,000 - 170,000
Member of Technical Staff — Data Infrastructure
Member of Technical Staff — Data Infrastructure

Causal Labs • San Francisco (CA)

On-site
USD 150,000 - 210,000
Data Infrastructure Engineer - Large-Scale AI Pipelines
Data Infrastructure Engineer - Large-Scale AI Pipelines

Genesis AI • Northern (KY)

Hybrid
USD 120,000 - 180,000
Member of Technical Staff — Data Infrastructure
Member of Technical Staff — Data Infrastructure

Kindredventures • San Francisco (CA)

On-site
USD 130,000 - 170,000
Data Infrastructure Engineer
Data Infrastructure Engineer

Samson Rose • San Francisco (CA)

On-site
USD 130,000 - 190,000
Data Infrastructure Engineer
Data Infrastructure Engineer

Pivot Robotics • San Francisco (CA)

On-site
USD 140,000 - 190,000
Senior Data Infrastructure Engineer, Petabyte Pipelines
Senior Data Infrastructure Engineer, Petabyte Pipelines

Genesis AI • United States

On-site
USD 140,000 - 210,000
Member of Technical Staff — Data Ingestion & Quality
Member of Technical Staff — Data Ingestion & Quality

Kindredventures • San Francisco (CA)

On-site
USD 120,000 - 180,000
Data Platform Engineer, Autonomy Analytics
Data Platform Engineer, Autonomy Analytics

FieldAI • Peoria (IL)

On-site
USD 110,000 - 160,000