Member of Technical Staff, Distributed Systems & Fleet

General Diffusion, Inc.

San Francisco, Northern (CA, KY)

Hybrid

USD 160,000 - 220,000

Full time

10 days ago
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Job summary

General Diffusion, Inc. is seeking a Member of Technical Staff to run and improve a multi-architecture testbed for research and experimentation.

You will own the fleet state, provisioning, and observability across heterogeneous hardware to ensure reproducible results. You will build dashboards and runbooks, drive incident response, and partner with teams to diagnose workload behavior on diverse silicon while maintaining reliability and auditable processes.

Qualifications

  • Experience operating production distributed systems on Linux, including deployment, diagnosis, and day-two operations.
  • Hands-on experience with bare-metal infrastructure and heterogeneous hardware lifecycles.
  • Strong observability mindset using metrics, logs, and traces to guide triage and analysis.

Responsibilities

  • Own inventory, provisioning, retirement, and readiness checks across the testbed; keep fleet state auditable.
  • Develop repeatable experiment lifecycles from reservation to return-to-service.
  • Capture context around real-world traces and failures to distinguish workload vs environment.
  • Build fleet-health signals, dashboards, and alerts connecting hosts, networks, accelerators, and experiments.
  • Maintain incident runbooks and lead post-incident improvements demonstrated through exercises.
  • Collaborate with data infrastructure on evidence contracts.

Skills

Distributed systems
Linux
Bare-metal infrastructure
Observability & monitoring
Incident response

Job description

Member of Technical Staff, Distributed Systems & Fleet

Keep a multi-architecture testbed reproducible, observable, and reliable.

Status Open

Area Infrastructure

Run the operational foundation for GD-X, General Diffusion’s bare-metal, multi-architecture testbed, so experiments on unlike silicon are repeatable and diagnosable rather than one-off machine events. You will make fleet state, experiment conditions, failures, and recovery visible to the researchers and systems teams learning how workloads behave across heterogeneous compute-without owning the learning policy, kernels, or placement implementation.

01 / The work

What you'll work on
  • Own machine inventory, bring-up, provisioning, retirement, and readiness checks across the heterogeneous testbed; make the state of each usable system explicit and auditable.
  • Build a repeatable experiment lifecycle-from reservation and environment preparation through execution, cleanup, and machine return-so comparable runs do not depend on undocumented operator steps.
  • Capture operational context around real-world traces and failures, including the machine and experiment conditions needed to distinguish fleet effects from workload behavior; partner with Measurement & Data Infrastructure on the shared evidence contract.
  • Develop fleet-health signals, dashboards, and alerts that connect host, network, accelerator, and experiment symptoms to actionable diagnosis while keeping paging paths understandable.
  • Engineer safe recovery paths for failed or suspect machines, including isolation, reprovisioning, validation, and documented return-to-service criteria.
  • Maintain practical incident runbooks and lead operational follow-through: triage, evidence preservation, root-cause analysis, and improvements validated through representative failure exercises.

02 / The background

What you bring
  • Experience operating production distributed systems on Linux, with ownership that spans deployment, diagnosis, and reliable day-two operation.
  • Hands-on experience with bare-metal infrastructure or a heterogeneous fleet, including hardware lifecycle, provisioning, configuration drift, and failure isolation.
  • Strong observability judgment: you can combine metrics, logs, traces, and targeted health checks into signals that support both rapid triage and retrospective analysis.
  • A rigorous approach to reproducibility in performance-sensitive or experimental systems, including recording the operational conditions that make results comparable.
  • Demonstrated incident-response discipline-clear runbooks, low-noise escalation, recovery procedures, and post-incident changes that are tested rather than merely documented.

03 / The evidence

What progress looks like
  • An auditable inventory and repeatable provisioning/readiness workflow provides evidence of each supported machine’s operational state before it enters an experiment.
  • Priority testbed runs yield comparable operational traces with sufficient machine and failure context to investigate whether observed differences arise from the workload or the environment, in collaboration with the data-infrastructure owner.
  • Dashboards, alerts, recovery runbooks, and recorded failure exercises demonstrate that representative fleet faults can be detected, contained, investigated, and returned to service through a defined operational path.

04 / In the system

Where this role fits

Owns physical-system evidence and operation, not RL algorithms or accelerator kernels.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Member of Technical Staff, Measurement & Data Infrastructure
Member of Technical Staff, Measurement & Data Infrastructure

General Diffusion, Inc. • San Francisco (CA), Northern (KY)

Hybrid
USD 180,000 - 240,000
Staff Engineer, Distributed Systems & Fleet Reliability
Staff Engineer, Distributed Systems & Fleet Reliability

General Diffusion, Inc. • San Francisco (CA), Northern (KY)

Hybrid
USD 160,000 - 220,000
Member of Technical Staff, GPU Systems & Fabric
Member of Technical Staff, GPU Systems & Fabric

General Diffusion, Inc. • San Francisco (CA), Northern (KY)

Hybrid
USD 180,000 - 250,000
Member of Technical Staff, Heterogeneous Runtime & Placement
Member of Technical Staff, Heterogeneous Runtime & Placement

General Diffusion, Inc. • San Francisco (CA), Northern (KY)

Hybrid
USD 170,000 - 250,000
Software Engineer — Fleet
Software Engineer — Fleet

Specter • San Francisco (CA)

On-site
USD 180,000 - 240,000
Member of Technical Staff, Compute World Models
Member of Technical Staff, Compute World Models

General Diffusion, Inc. • San Francisco (CA), Northern (KY)

Hybrid
USD 180,000 - 240,000
Member of Technical Staff, RL & Compute Environments
Member of Technical Staff, RL & Compute Environments

General Diffusion, Inc. • San Francisco (CA), Northern (KY)

Hybrid
USD 180,000 - 280,000
Member of Technical Staff, Kernels
Member of Technical Staff, Kernels

General Diffusion, Inc. • San Francisco (CA), Northern (KY)

Hybrid
USD 180,000 - 280,000
Member of Technical Staff — Training Infrastructure
Member of Technical Staff — Training Infrastructure

Human Intuition Inc. • New York (NY)

On-site
USD 150,000 - 210,000
Software Engineer, Machine Lifecycle
Software Engineer, Machine Lifecycle

Career Techniques • New York (NY)

On-site
USD 150,000 - 250,000