Senior Data Infrastructure Engineer

Greylock Partners

San Francisco (CA)

On-site

USD 180,000 - 260,000

Full time

5 days ago
Be an early applicant
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Job summary

Greylock Partners is partnering with an early-stage AI company to build a durable data foundation for large-scale video and multimodal data. We seek a senior engineer who has owned petabyte-scale storage and data platforms in production, shaping architectures with long-term impact.

You'll design storage, processing pipelines, and observability, collaborating with founders to turn early product needs into scalable, robust data infrastructure for researchers and agents.

Qualifications

  • Direct ownership of petabyte-scale storage or data platforms in production.
  • Strong distributed-systems judgment around storage, throughput, reliability, and cost.
  • Experience building data systems for demanding workloads like ML training or large-scale analytics.
  • Ability to investigate failures across system boundaries and improve design.

Responsibilities

  • Design and build storage and processing systems for large volumes of video and data.
  • Develop reliable pipelines for high-throughput multimodal processing.
  • Make deliberate choices about data layout, lifecycle, and cost as volume grows.
  • Build infrastructure linking captured activity to research experiments and production workflows.
  • Establish observability and operating practices for data quality and performance.
  • Work with founders to shape durable data architecture.

Skills

Petabyte-scale storage
Distributed systems
Data platforms
ML data infrastructure

Job description

Greylock is partnering with an early-stage AI company focused on a gap between model capability and adoption. AI can already change how work gets done, but bringing agents into a real organization remains slow, manual, and difficult to scale beyond individual pilots.

The company is building a repeatable path from an initial AI-agent deployment to broader organizational impact. Its goal is to make those agents correct, trustworthy, robust, and scalable-and to help companies build on each successful deployment.

Summary

We're looking for a senior engineer who has personally built and operated petabyte-scale storage or data infrastructure. You'll help design the foundation for a product whose data volume could eventually reach petabytes per day.

The immediate challenge is large-scale storage and processing, including the pipelines that turn raw video into useful representations. This is less about low-latency streaming than about building a dependable, efficient system that can ingest large volumes, process them at high throughput, and make the results available to researchers and agents.

You'll work directly with the founders and have unusual influence over the architecture from the beginning.

What You’ll Own
  • Design and build storage and processing systems for large volumes of video, accessibility data, and derived workflow data
  • Develop reliable pipelines for high-throughput multimodal captioning and other video-processing workloads
  • Make deliberate choices about data layout, lifecycle, retention, retrieval, reliability, and cost as volume grows
  • Build the infrastructure that connects captured activity to research experiments and production agent workflows
  • Establish observability and operating practices that make data quality, failures, and system performance visible
  • Work with the founders to turn early product needs into a durable data architecture
What We’re Looking For
  • Direct ownership of petabyte-scale storage or data platforms in production, including architectural decisions and operational responsibility
  • Strong distributed-systems judgment around storage, throughput, reliability, and cost
  • Experience building data systems that serve demanding downstream workloads, such as ML training, multimodal processing, or large-scale analytics
  • The ability to investigate failures across system boundaries and improve the underlying design
  • Comfort building hands-on in a small team where requirements and scale will evolve quickly

Experience with video or multimodal data, object storage, lakehouse systems, or ML data infrastructure is valuable. We care most about evidence that you have owned the storage problem at scale, rather than only run pipelines over large datasets. Real-time streaming expertise is useful, but it is not the center of this role.

About Us

Greylock is an early-stage investor in companies including Airbnb, LinkedIn, Dropbox, Workday, Cloudera, Facebook, Instagram, Roblox, Coinbase, and Palo Alto Networks. Learn more at greylock.com.

How We Work

Greylock's Core Talent team provides free candidate referrals and introductions to our active investments. This posting is for direct employment with one of our portfolio companies. We review every application and reach out directly when we believe there's a strong potential fit.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Data Infra Engineer - Petabyte-Scale Storage & Video Pipelines
Data Infra Engineer - Petabyte-Scale Storage & Video Pipelines

Greylock Partners • San Francisco (CA)

On-site
USD 180,000 - 260,000
Founding Data Engineer (Product/Finance/GTM)
Founding Data Engineer (Product/Finance/GTM)

Greylock Partners • San Francisco (CA)

On-site
USD 180,000 - 240,000
AI Researcher, Agent Systems
AI Researcher, Agent Systems

Greylock Partners • San Francisco (CA)

On-site
USD 150,000 - 210,000
Research Member of Technical Staff- Data Infrastructure
Research Member of Technical Staff- Data Infrastructure

Rhoda AI • Palo Alto (CA)

On-site
USD 180,000 - 280,000
Research Member of Technical Staff- Data Infrastructure
Research Member of Technical Staff- Data Infrastructure

Rhoda AI • Mountain View (CA)

On-site
USD 180,000 - 240,000
Founding Engineer — AI Agent Security
Founding Engineer — AI Agent Security

Greylock Partners • San Francisco (CA)

On-site
USD 180,000 - 280,000
Senior Software Engineer, Backend
Senior Software Engineer, Backend

Mecka AI • New York (NY)

On-site
USD 120,000 - 150,000
Engineering Manager (Backend/Infra)
Engineering Manager (Backend/Infra)

Greylock Partners • San Francisco (CA)

On-site
USD 240,000 - 320,000
Member of Technical Staff - Storage Infrastructure
Member of Technical Staff - Storage Infrastructure

AI Chopping Block • San Francisco (CA), Northern (KY)

On-site
USD 150,000 - 300,000
Member of Technical Staff - Storage Infrastructure
Member of Technical Staff - Storage Infrastructure

Prime Intellect AI • San Francisco (CA)

On-site
USD 150,000 - 300,000