Member of Technical Staff – Fleet Operations

Hamilton Barnes ?

San Francisco (CA)

On-site

USD 190,000 - 250,000

Full time

14 days+
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Job summary

stealth neocloud is seeking a senior infrastructure leader to own the fleet control plane and full machine lifecycle for datacenter-scale AI workloads. You will drive provisioning, health, and lifecycle across Apple hardware generations, collaborating with multiple teams to ensure safe, observable operations.

You will build distributed systems and champion agentic development, defining health signals, SLOs, and recovery strategies while guiding incident response in a high-scale environment.

Qualifications

  • Built and operated large production systems that other teams depend on.
  • Strong software engineering skills in Go, Python, Rust or similar languages.
  • Experience with distributed systems, control planes, state machines, controllers, or durable workflows.
  • Strong knowledge of Linux, macOS, or Unix systems, with comfort across boot flows, processes, networking, storage, containers, and system performance.
  • Experience with bare-metal compute, machine provisioning, Kubernetes, workload schedulers, or large server fleets.
  • Experience connecting node-level software to distributed control planes or automated operators.
  • A track record of leading complex, multi-team infrastructure work from strategy through production.
  • Strong operational judgment, with a focus on designing for partial failure, safe retries, auditability, and recovery.

Responsibilities

  • Own the technical strategy and roadmap for the fleet control plane and full machine lifecycle, leading complex work across teams and systems.
  • Build the distributed control plane and node-level software that manage inventory, configuration, health, and lifecycle state — ensuring every action is safe, observable, auditable, and recoverable.
  • Automate capacity ingestion across Apple hardware generations, including provisioning, validation, configuration, updates, reimaging, diagnostics, repair, and return to service.
  • Connect fleet health and capacity to workload scheduling to improve availability, placement, utilization, recovery time, and time-to-production for new capacity.
  • Champion agentic development and operations — using coding agents throughout investigation, implementation, testing, and operations, and building interfaces that let software agents inspect state, take safe action, verify results, and elevate exceptions.
  • Establish strong operational practices: define health signals and service objectives, lead incidents, improve on-call health, and convert failures into lasting system improvements.

Skills

Go
Python
Distributed systems
Rust

Tools

Kubernetes
Linux
macOS

Job description

Join a stealth neocloud that makes Apple hardware (macOS and Apple Silicon) available at datacenter scale for AI workloads - powering computer-use models, long-running agents, and agentic engineering for their customers.

The Fleet team owns the software layer that turns raw compute into a reliable, elastic fleet:

  • Defines how machines enter, operate, recover, and leave the fleet
  • Connects physical capacity to workload demand
  • Maximizes healthy, schedulable capacity while minimizing capacity lost to provisioning failures, hardware faults, or incomplete recovery
  • Serves as a proving ground for agentic engineering — building systems that software agents can safely inspect and operate, so agents handle routine investigation, change, and recovery while humans set policy, manage risk, and resolve novel failures
  • Works across hardware, operating systems, networking, scheduling, security, and data center operations

Key Responsibilities

  • Own the technical strategy and roadmap for the fleet control plane and full machine lifecycle, leading complex work across teams and systems
  • Build the distributed control plane and node-level software that manage inventory, configuration, health, and lifecycle state — ensuring every action is safe, observable, auditable, and recoverable
  • Automate capacity ingestion across Apple hardware generations, including provisioning, validation, configuration, updates, reimaging, diagnostics, repair, and return to service
  • Connect fleet health and capacity to workload scheduling to improve availability, placement, utilization, recovery time, and time-to-production for new capacity
  • Champion agentic development and operations — using coding agents throughout investigation, implementation, testing, and operations, and building interfaces that let software agents inspect state, take safe action, verify results, and elevate exceptions
  • Establish strong operational practices: define health signals and service objectives, lead incidents, improve on-call health, and convert failures into lasting system improvements

Requirements

  • Built and operated large production systems that other teams depend on
  • Strong software engineering skills in Go, Python, Rust, or a similar language
  • Experience with distributed systems, control planes, state machines, controllers, or durable workflows
  • Strong knowledge of Linux, macOS, or Unix systems, with comfort across boot flows, processes, networking, storage, containers, and system performance
  • Experience with bare-metal compute, machine provisioning, Kubernetes, workload schedulers, or large server fleets
  • Experience connecting node-level software to distributed control planes or automated operators
  • Deep experience using coding agents to build production software, including how to provide the context, tools, tests, and constraints required for reliable results
  • A track record of leading complex, multi-team infrastructure work from strategy through production
  • Strong operational judgment, with a focus on designing for partial failure, safe retries, auditability, and recovery

Travel and Work Authorization

  • The role requires the employee to work in the US with occasional travel to data centers in San Francisco Bay area.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Member of Technical Staff, Fleet
Member of Technical Staff, Fleet

Mount Thor • San Francisco (CA)

On-site
USD 180,000 - 240,000
Member of Technical Staff, Fleet
Member of Technical Staff, Fleet

Mount Thor, Inc. • San Francisco (CA), Northern (KY)

Hybrid
USD 160,000 - 210,000
Staff Engineer, Fleet Control Plane & AI Compute
Staff Engineer, Fleet Control Plane & AI Compute

Hamilton Barnes ? • San Francisco (CA)

On-site
USD 190,000 - 250,000
Member of Technical Staff, Fleet Operations
Member of Technical Staff, Fleet Operations

Mount Thor, Inc. • San Francisco (CA), Northern (KY)

Hybrid
USD 180,000 - 240,000
Software Engineer, Infrastructure Services (Cloud Network Fault Tolerance)
Software Engineer, Infrastructure Services (Cloud Network Fault Tolerance)

JobCubby • California (MO), Northern (KY)

Hybrid
USD 150,000 - 210,000
Full Stack Software Engineer - ML Compute Capacity
Full Stack Software Engineer - ML Compute Capacity

Apple Inc. • Santa Clara (CA)

On-site
USD 184,700 - 324,800
Apple benefits
Stock programs
Relocation assistance
+1
Big Data Systems Engineer
Big Data Systems Engineer

Socket.dev • Austin (TX)

On-site
USD 140,000 - 190,000
Software Engineer — Fleet
Software Engineer — Fleet

Specter • San Francisco (CA)

On-site
USD 180,000 - 240,000
Operations Software Engineer- Customer Systems
Operations Software Engineer- Customer Systems

Socket.dev • Austin (TX)

On-site
USD 140,000 - 220,000
Sr DevOps Engineer, Retail and Marcom Engineering
Sr DevOps Engineer, Retail and Marcom Engineering

Socket.dev • Austin (TX)

On-site
USD 150,000 - 230,000