Research Engineer

Pantograph

San Francisco (CA)

On-site

USD 180,000 - 230,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Pantograph is training general models that start by watching internet-scale video and end up on robots. We think the path to capable robots runs through general intelligence rather than narrow, robot-specific skills.

We're scaling simple methods across video, real-world data, and our fleet of durable robots. You’ll work across the boundary between research and engineering: implementing new ideas, scaling experiments across large GPU clusters, building the systems that let us iterate quickly, and

Qualifications

  • Experience training models on large GPU clusters.
  • Proficiency with Kubernetes and distributed systems.
  • Ability to scale experiments across multiple clusters and datasets.
  • Track record building production-like ML systems with observable metrics.

Responsibilities

  • Design scalable training pipelines and multimodal models.
  • Scale experiments across large GPU clusters and data stores.
  • Build systems enabling rapid iteration between research and production.
  • Analyze failures, improve observability, and drive system reliability.
  • Collaborate across research and engineering to deploy capabilities.

Skills

Kubernetes
Distributed systems
Large-scale data processing
Multimodal model training
Experimentation & debugging
Research-to-product bridge
Python / ML engineering

Tools

JAX
CUDA kernels
Linux kernel programming
GPU clusters
TensorFlow / PyTorch

Job description

Pantograph is training general models that start by watching internet-scale video and end up on robots. We think the path to capable robots runs through general intelligence rather than narrow, robot-specific skills. We're scaling simple methods across video games, real-world video, and our own fleet of affordable, durable robots.

We're looking for a research engineer to help us train increasingly capable models across enormous and diverse datasets.

You’ll work across the boundary between research and engineering: implementing new ideas, scaling experiments across large GPU clusters, building the systems that let us iterate quickly, and figuring out why things aren’t working. The work spans large-scale model training, multimodal representation learning, reinforcement learning, data processing, evaluation, and the infrastructure required to support all of it.

You might be a good fit if you:

  • Have trained models across large GPU clusters and are comfortable working with Kubernetes

  • Have built or operated complex distributed systems

  • Have worked with multi-terabyte or multi-petabyte datasets

  • Are comfortable with large-scale data processing tools

  • Care deeply about observability and collect enough metrics to understand what every part of a system is doing

  • Are comfortable moving between research code and production-quality systems

  • Like running experiments, getting surprising results, and digging in until you understand why

  • Move quickly and reach for simple approaches before complicated ones

Nice to have:

  • Experience with JAX

  • Experience writing CUDA kernels or otherwise optimizing GPU workloads

  • Low-level Linux or kernel programming experience

  • Experience with large-scale video or multimodal datasets

  • Experience building training or evaluation infrastructure

  • Experience with distributed training

  • Experience deploying models into real-world systems, especially robotics

We care much more about what you've built than any specific credential. We're a small, fast-moving team working together in person in San Francisco. If you're excited about architecting novel systems at unprecedented scale, we'd love to talk.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Research Engineer: Scale-Driven AI Systems & Robotics
Research Engineer: Scale-Driven AI Systems & Robotics

Pantograph • San Francisco (CA)

On-site
USD 180,000 - 230,000
Research Engineer
Research Engineer

Harnham • United States

On-site
USD 120,000 - 150,000
Research Engineer, Infrastructure, Training Systems
Research Engineer, Infrastructure, Training Systems

Thinkingmachines • San Francisco (CA)

On-site
USD 350,000 - 475,000
Generous health, dental, and vision benefits
Unlimited PTO
Paid parental leave
+1
Research Member of Technical Staff- Training Systems
Research Member of Technical Staff- Training Systems

Rhoda AI • Mountain View (CA)

On-site
USD 140,000 - 180,000
Research Member of Technical Staff- Training Systems
Research Member of Technical Staff- Training Systems

Rhoda AI • Mountain View (CA)

On-site
USD 150,000 - 200,000
Research Engineer, ML Infrastructure
Research Engineer, ML Infrastructure

cognition • San Francisco (CA)

On-site
USD 180,000 - 250,000
Research Engineer Infrastructure Training Systems
Research Engineer Infrastructure Training Systems

Thinking Machines Lab • San Francisco (CA)

On-site
USD 350,000 - 475,000
Generous health benefits
Unlimited PTO
Paid parental leave
+1
Research Scientist: Pretraining
Research Scientist: Pretraining

Generalist • San Francisco (CA)

On-site
USD 120,000 - 150,000
Research Member of Technical Staff- Training Systems
Research Member of Technical Staff- Training Systems

Rhoda AI • Palo Alto (CA)

On-site
USD 210,000 - 320,000
Research Member of Technical Staff - Training Platform
Research Member of Technical Staff - Training Platform

Rhoda AI • Mountain View (CA)

On-site
USD 100,000 - 140,000