Member of Technical Staff, Software Engineer (Data Infrastructure)

Orbifold AI, Inc.

Palo Alto (CA)

On-site

USD 140,000 - 200,000

Full time

14 days+
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Job summary

engineeringjobs.net, Inc. seeks a data infrastructure lead to design and own the platform for multimodal capture, covering ingest, storage, indexing, and delivery.

You will ensure durability, queryability, and performance across massive datasets for large-scale distributed training runs. The role requires 3+ years of production-grade data infra experience, strong Python and systems programming skills, and deep expertise in storage layout and distributed processing frameworks.

Qualifications

  • At least 3 years of experience building production-grade data infrastructure with strong proficiency in Python and systems programming.
  • Deep expertise in storage layout, distributed processing frameworks, and data reproducibility at scale.

Responsibilities

  • Design and own the data infrastructure platform for multimodal capture, including ingest, storage, indexing, and delivery.
  • Ensure massive datasets are durable, queryable, and performant enough to support large-scale distributed training runs.

Skills

Python
Systems programming
Data infrastructure
Storage architecture
Distributed processing
Ray
Spark
Data indexing
Data versioning
Video processing
Sensor data
Schema evolution
Performance optimization
Data lineage
Distributed training

Job description

You will design and own the data infrastructure platform for multimodal capture, including ingest, storage, indexing, and delivery. Your work ensures that massive datasets are durable, queryable, and performant enough to support large-scale distributed training runs.

Requirements: Candidates must have at least 3 years of experience building production-grade data infrastructure with strong proficiency in Python and systems programming. You should possess deep expertise in storage layout, distributed processing frameworks, and data reproducibility at scale.

Key Skills: Data infrastructure, Python, Systems programming, Storage architecture, Distributed processing, Ray, Spark, Data indexing, Data versioning, Video processing, Sensor data, Schema evolution, Performance optimization, Data lineage, Distributed training

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Data Infrastructure Engineer for Scalable Multimodal Data
Data Infrastructure Engineer for Scalable Multimodal Data

Orbifold AI, Inc. • Palo Alto (CA)

On-site
USD 140,000 - 200,000
Software Engineer (Ray Data)
Software Engineer (Ray Data)

EngRadar • San Francisco (CA)

On-site
USD 140,000 - 190,000
Member of Technical Staff — Data Infrastructure
Member of Technical Staff — Data Infrastructure

Kindredventures • San Francisco (CA)

On-site
USD 130,000 - 170,000
Software Engineer, Data Flywheel Platform
Software Engineer, Data Flywheel Platform

Necessary Ventures • Sunnyvale (CA)

On-site
USD 140,000 - 190,000
Tech Lead, Data & Inference Engineer
Tech Lead, Data & Inference Engineer

Catalyst Labs • Connecticut

On-site
USD 120,000 - 180,000
Tech Lead, Data & Inference Engineer
Tech Lead, Data & Inference Engineer

Catalyst Labs • New York (NY)

On-site
USD 120,000 - 160,000
Member of Technical Staff — Data Ingestion & Quality
Member of Technical Staff — Data Ingestion & Quality

Kindredventures • San Francisco (CA)

On-site
USD 120,000 - 180,000
Tech Lead, Data & Inference Engineer
Tech Lead, Data & Inference Engineer

Catalyst Labs • Seattle (WA)

On-site
USD 120,000 - 160,000
Data Infrastructure
Data Infrastructure

Genesis AI • United States

On-site
USD 140,000 - 210,000
Tech Lead, Data & Inference Engineer
Tech Lead, Data & Inference Engineer

Catalyst Labs • Jacksonville (FL)

On-site
USD 130,000 - 160,000