Infrastructure & Systems Lead

Pantograph

San Francisco (CA)

On-site

USD 180,000 - 240,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Pantograph is building a scalable infrastructure stack for a fleet of robots, streaming large video data to training clusters, with frequent model updates and strict real-time latency goals.

The role requires owning the end-to-end system, including networking, storage, databases, embedded software, and deployment pipelines, and optimizing GPU performance across the stack.

Qualifications

  • Designed and operated large-scale distributed systems from scratch.
  • Managed fleets of hundreds or thousands of computers.
  • Deep experience with GPU performance optimization, including CUDA kernels.
  • Worked on real-time embedded systems or robotics infrastructure.
  • Built petabyte-scale storage and database systems.
  • Experience with high-performance networking and video encoding pipelines.

Responsibilities

  • Architect and own the infrastructure pipeline behind a fleet of robots and inference clusters.
  • Touch networking, storage, databases, embedded software, deployment systems, and GPU optimization.
  • Own the architecture decisions that tie these components together to meet real-time latency budgets.

Skills

Distributed systems design
Fleet management
GPU performance optimization
Real-time embedded systems
Petabyte-scale storage
High-performance networking
Video encoding pipelines

Tools

CUDA kernels

Job description

Pantograph is training general models that start by watching internet-scale video and end up on robots. We think the path to capable robots runs through general intelligence rather than narrow, robot-specific skills. We're scaling simple methods across video games, real-world video, and our own fleet of affordable, durable robots.

We're looking for someone to architect and own the entire infrastructure pipeline behind that fleet: thousands of robots with embedded GPUs, communicating over wifi to inference clusters, streaming tens of petabytes of video to training clusters, with new model weights deployed every few minutes — all operating within strict real-time latency budgets.

This role requires someone who can hold an entire system in their head and optimize it end-to-end. You'll touch networking, storage, databases, embedded software, deployment systems, and GPU optimization — and you'll own the architecture decisions that tie them together.

You might be a good fit if you have:

  • Designed and operated large-scale distributed systems from scratch

  • Managed fleets of hundreds or thousands of computers

  • Deep experience with GPU performance optimization, including writing custom CUDA kernels

  • Worked on real-time embedded systems or robotics infrastructure

  • Built petabyte-scale storage and database systems

  • Experience with high-performance networking and video encoding pipelines

Nice to have:

  • Rust and low-level performance optimization experience

  • Experience taking infrastructure from prototype to production in a small team

We care much more about what you've built than any specific credential. We're a small, fast-moving team working together in person in San Francisco. If you're excited about architecting novel systems at unprecedented scale, we'd love to talk.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Research Engineer
Research Engineer

Pantograph • San Francisco (CA)

On-site
USD 180,000 - 230,000
Software Engineer
Software Engineer

Pantograph • San Francisco (CA)

On-site
USD 140,000 - 210,000
Software Engineer, Cloud Infrastructure
Software Engineer, Cloud Infrastructure

Socket.dev • San Francisco (CA)

On-site
USD 140,000 - 210,000
Staff Software Engineer (AI Infrastructure)
Staff Software Engineer (AI Infrastructure)

DeepRec.ai • Palo Alto (CA)

On-site
USD 180,000 - 320,000
Mechanical Design Engineer
Mechanical Design Engineer

Pantograph PBC • San Francisco (CA)

On-site
USD 120,000 - 160,000
Cluster Design
Cluster Design

Blue Signal Search • San Francisco (CA)

On-site
USD 150,000 - 230,000
AI Infrastructure Lead
AI Infrastructure Lead

Outsourceit • San Francisco (CA)

On-site
USD 120,000 - 170,000
AI Network Performance and Reliability Engineer
AI Network Performance and Reliability Engineer

AMP PBC • San Francisco (CA)

On-site
USD 180,000 - 240,000
Member of Technical Staff – Software Engineer, GPU Cluster Infrastructure
Member of Technical Staff – Software Engineer, GPU Cluster Infrastructure

Perplexity • San Francisco (CA)

On-site
USD 180,000 - 240,000
Member of Technical Staff - GPU Infrastructure
Member of Technical Staff - GPU Infrastructure

Prime Intellect • United States

On-site
USD 120,000 - 150,000