ML Infra Engineer, Platform

Physical Intelligence

San Francisco (CA)

On-site

USD 180,000 - 210,000

Full time

2 days ago
Be an early applicant
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Job summary

Physical Intelligence in San Francisco is seeking a seasoned Platform Infrastructure Engineer to own and scale AI-native infrastructure. You will operate Kubernetes clusters, build a scalable microservice platform, and drive security, observability, and cost-aware design across multi-cloud environments.

You will collaborate with researchers and engineers to translate fast-moving needs into reusable infrastructure, owning systems end-to-end from design to operation.

Qualifications

  • Strong first-principles thinking and ownership.
  • Experience with cloud platforms (GCP, AWS) and distributed systems.
  • In-depth experience with infrastructure-as-code and containerization.

Responsibilities

  • Own and scale AI-native infrastructure: operate and evolve Kubernetes clusters and service deployment patterns, support safe rollouts, upgrades, and rollback strategies.
  • Drive security, observability and cost-aware infrastructure, surface reliability and performance issues early, improve cost visibility.
  • Harden platform foundations across multi-cloud, design authentication/authorization flows, networking, quota and rate-limiting services.
  • Improve developer experience with clear interfaces and self-serve infrastructure usage.
  • Collaborate with researchers and engineers to own systems end-to-end from design to operation.

Skills

First-principles thinking
Agent infrastructure
Cloud platforms (GCP, AWS)
Distributed systems
Kubernetes
Observability
Security
Sandboxing
Cost optimization
Cross-functional communication

Tools

Kubernetes
Terraform
Docker
GCP
AWS
CI/CD

Job description

Who We Are

Physical Intelligence is bringing general-purpose AI into the physical world. We are a team of engineers, scientists, roboticists, and company builders developing foundation models and learning algorithms to power the robots of today and the physically-actuated devices of the future.


The Team

The Infrastructure team builds and operates the backbone of everything PI does: from training state-of-the-art VLA models, to orchestrating large-scale simulation, to reliably deploying intelligence across fleets of physical robots. The team works closely with researchers, robotics runtime, product, and platform engineers to ensure infrastructure scales from prototype to production-grade deployments.


In This Role You Will


  • Own and scale AI-native infrastructure: You will operate and evolve Kubernetes clusters and service deployment patterns, and help build a scalable microservice platform for internal systems such as evaluation services, operational tooling, and internal APIs with agent-use as the primary interface. This includes supporting safe rollouts, upgrades, and rollback strategies.


  • Drive security, observability and cost-aware infrastructure: You will treat security, logging, metrics, tracing, and alerting as first-class platform primitives, and build systems that surface reliability and performance issues early. You will also help improve cost visibility and enable cost-aware decision-making at the infrastructure level.


  • Harden platform foundations: You will own core infrastructure with multi-cloud considerations, designing authentication/authorization flows, networking architecture, quota and rate-limiting services, and cloud primitives that behave predictably. A major part of this work is reducing infra churn by standardizing patterns, abstractions, and interfaces.


  • Improve developer experience: You will build clear, documented interfaces for using platform infrastructure, reducing the gap between “I need infra” and “I can run my workload.” This includes supporting consistent local vs. remote development workflows and improving self‑serve infrastructure usage.


  • Collaborate and lead through ownership: You will work closely with researchers and other engineers to understand requirements and constraints, translate fast-moving needs into reusable infrastructure, and own systems end-to-end, from design through operation.



What We Hope You'll Bring


  • Extremely strong first-principles thinking.


  • Familiarity with agent infrastructure, observability, security and sandboxing


  • Deep experience with cloud platforms (GCP, AWS) and distributed systems: compute orchestration, networking, autoscaling, service meshes, load balancing.


  • Ability to reason about system bottlenecks, performance tuning, and cost optimizations across compute, networking, and storage.


  • Comfort with Kubernetes, cluster-level reliability, and service-oriented architectures.


  • Solid intuition around scalability, performance, and failure modes.


  • Experience with infrastructure-as-code (e.g. Terraform), containerization, and modern platform engineering practices.


  • Familiarity with logging, metrics, tracing, incident response, SLOs, and debugging complex distributed systems.


  • Strong cross‑functional communication and ownership mindset.


  • Experience (4-6 years) working in fast‑moving or early‑stage environments where ambiguity is normal with demonstrated growth trajectory.



Ultimately, we’re looking for someone who can spin up quickly on unfamiliar and ambiguous domains, has a strong sense of ownership, and cares deeply about our mission. Even if you don’t check all the boxes above, we strongly encourage you to apply if this sounds like you.


Bonus Points If You Have


  • Experience with large‑scale ML training, evaluation, or simulation infrastructure.


  • Experience with secrets management systems (e.g., Doppler).


  • Background in observability, cost optimization, or internal platform tooling.


  • Exposure to robotics, simulation, or real‑time systems.



Pursuant to the San Francisco Fair Chance Ordinance, we will consider for employment qualified applicants with arrest and conviction records.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

ML Infra Engineer (Data Systems)
ML Infra Engineer (Data Systems)

Physical Intelligence • San Francisco (CA)

On-site
USD 180,000 - 240,000
ML Infra Engineer, Modeling
ML Infra Engineer, Modeling

Physical Intelligence • San Francisco (CA)

On-site
USD 180,000 - 240,000
ML engineer - API Platform
ML engineer - API Platform

Physical Intelligence • San Francisco (CA)

On-site
USD 150,000 - 210,000
Software Engineer, AI Productivity
Software Engineer, AI Productivity

Physical Intelligence • San Francisco (CA)

On-site
USD 140,000 - 230,000
Fullstack Software Engineer, Robot Interfaces
Fullstack Software Engineer, Robot Interfaces

Physical Intelligence • San Francisco (CA)

On-site
USD 140,000 - 210,000
Staff Platform Engineer - Developer Infrastructure
Staff Platform Engineer - Developer Infrastructure

Persona AI, Inc. • Houston (TX), Northern (KY)

On-site
USD 150,000 - 210,000
Competitive compensation
Performance-based bonus
Medical benefits (99% employer covered
+3
Software Engineer, AI Productivity
Software Engineer, AI Productivity

Precision Labs • San Francisco (CA), Northern (KY)

Hybrid
USD 140,000 - 190,000
ML engineer - API Platform Physical Intelligence · San Francisco, CA Full-time · On-site — 3 hours ago
ML engineer - API Platform Physical Intelligence · San Francisco, CA Full-time · On-site — 3 hours ago

Emploive • San Francisco (CA)

On-site
USD 180,000 - 240,000
Member of Technical Staff — Infrastructure
Member of Technical Staff — Infrastructure

Observable Intuition, Inc. • New York (NY), Northern (KY)

Hybrid
USD 120,000 - 200,000
Member of Technical Staff — Infrastructure
Member of Technical Staff — Infrastructure

Collective Intuition, Inc. • New York (NY), Northern (KY)

Hybrid
USD 120,000 - 200,000