Member of Technical Staff, RL & Compute Environments

General Diffusion, Inc.

San Francisco, Northern (CA, KY)

Hybrid

USD 180,000 - 280,000

Full time

10 days ago
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Job summary

General Diffusion, Inc. in San Francisco seeks a Member of Technical Staff to turn compute predictions into constrained, measurable decisions.

You will define environment abstractions, evaluation contracts, and baselines to guide placement and resource allocation without owning production paths. You will build offline evaluation pipelines, report calibration limits and failure cases, and collaborate with ML researchers, runtime engineers, data teams, and safety reviewers to ensure rigorous,

Qualifications

  • Experience with RL, contextual bandits, planning, or off-policy evaluation where state distributions change and decisions have delayed system effects.
  • Ability to turn a systems problem into a falsifiable experimental environment, including action semantics, reward trade-offs, baselines, held-out conditions, and variance-aware measurement.
  • Strong programming and experimental-systems practice: build reliable research pipelines, inspect failures, and make experiments repeatable.
  • Practical scheduling or resource-allocation intuition, including trade-offs among latency, throughput, utilization, contention, and feasibility across heterogeneous resources.
  • Sound judgment about safe empirical work: distinguish simulation or offline evidence from online evidence, set bounded experiment conditions, and state uncertainty plainly.

Responsibilities

  • Design representative compute-environment abstractions: state, feasible placement actions, transition assumptions, rewards, and explicit constraint signals for heterogeneous workloads.
  • Build reproducible offline evaluation and counterfactual-analysis workflows that compare candidate policies against baselines using held-out workload, architecture, and drift conditions.
  • Develop policy-learning and planning experiments that use action-conditioned performance predictions while reporting calibration limits and uncertainty-sensitive failure cases.
  • Define shadow-evaluation and narrowly scoped online-experiment protocols with predeclared metrics, stop conditions, and rollback criteria; partner with Runtime on execution but do not own the production placement path.
  • Measure how policy quality changes with workload mix, hardware capability differences, delayed effects, and distribution shift, separating reward improvement from operational reliability.
  • Communicate decision evidence and known limitations to World Models, Runtime, and Safety & Verification; independent controls—not this role’s reward or policy—determine whether an action is permitted.

Skills

Reinforcement learning
Contextual bandits
Planning
Off-policy evaluation
Experiment design

Job description

Member of Technical Staff, RL & Compute Environments

Turn compute predictions into constrained, measurable decisions without delegating safety to a reward function.

Status Open

Area Research

Build the decision-learning environments that turn General Diffusion’s predictions of heterogeneous compute behavior into measurable placement and resource-allocation choices. You will define constrained RL problems, make offline evidence credible enough to inform bounded shadow or online evaluation, and keep the policy’s optimization work distinct from the independent controls that authorize actions. The work connects compute world models to policy optimization across unlike hardware without owning production execution, telemetry infrastructure, or safety enforcement.

01 / The work
What you’ll work on
  • Design representative compute-environment abstractions: state, feasible placement actions, transition assumptions, rewards, and explicit constraint signals for heterogeneous workloads.
  • Build reproducible offline evaluation and counterfactual-analysis workflows that compare candidate policies against baselines using held-out workload, architecture, and drift conditions.
  • Develop policy-learning and planning experiments that use action-conditioned performance predictions while reporting calibration limits and uncertainty-sensitive failure cases.
  • Define shadow-evaluation and narrowly scoped online-experiment protocols with predeclared metrics, stop conditions, and rollback criteria; partner with Runtime on execution but do not own the production placement path.
  • Measure how policy quality changes with workload mix, hardware capability differences, delayed effects, and distribution shift, separating reward improvement from operational reliability.
  • Work with Measurement & Data Infrastructure to specify the environment/action/outcome evidence needed for reproducible training and evaluation slices, without taking ownership of the underlying data platform.
  • Communicate decision evidence and known limitations to World Models, Runtime, and Safety & Verification; independent controls—not this role’s reward or policy—determine whether an action is permitted.
02 / The background
What you bring
  • Demonstrated work in reinforcement learning, contextual bandits, planning, or off-policy evaluation where state distributions change and decisions have delayed system effects.
  • Ability to turn a systems problem into a falsifiable experimental environment, including action semantics, reward trade-offs, baselines, held-out conditions, and variance-aware measurement.
  • Strong programming and experimental-systems practice: build reliable research pipelines, inspect failures, and make experiments repeatable rather than optimizing only a headline metric.
  • Practical scheduling or resource-allocation intuition, including the trade-offs among latency, throughput, utilization, contention, and feasibility across heterogeneous resources.
  • Sound judgment about safe empirical work: distinguish simulation or offline evidence from online evidence, set bounded experiment conditions, and state uncertainty plainly.
  • Clear cross-functional communication with ML researchers, runtime engineers, data/measurement partners, and independent safety reviewers.
03 / The evidence
What progress looks like
  • A documented compute-environment and evaluation contract exists for priority placement decisions, with explicit action feasibility, reward trade-offs, constraint signals, baselines, and known modeling assumptions.
  • Candidate policies can be compared reproducibly on held-out workload and hardware conditions, with decision-quality, variance, drift, and failure-mode results that make limits visible rather than obscuring them in aggregate rewards.
  • For experiments that progress beyond offline evaluation, evidence packages define predeclared success measures, bounded authority, stop/rollback conditions, and results from shadow or other appropriately controlled evaluation; Safety & Verification retains permissioning authority.
04 / In the system
Where this role fits

Owns policy learning and evidence; independent controls decide whether an action is permitted.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Member of Technical Staff, Compute World Models
Member of Technical Staff, Compute World Models

General Diffusion, Inc. • San Francisco (CA), Northern (KY)

Hybrid
USD 180,000 - 240,000
Member of Technical Staff, Safety & Formal Verification
Member of Technical Staff, Safety & Formal Verification

General Diffusion, Inc. • San Francisco (CA), Northern (KY)

Hybrid
USD 180,000 - 260,000
Member of Technical Staff, Heterogeneous Runtime & Placement
Member of Technical Staff, Heterogeneous Runtime & Placement

General Diffusion, Inc. • San Francisco (CA), Northern (KY)

Hybrid
USD 170,000 - 250,000
Member of Technical Staff, Measurement & Data Infrastructure
Member of Technical Staff, Measurement & Data Infrastructure

General Diffusion, Inc. • San Francisco (CA), Northern (KY)

Hybrid
USD 180,000 - 240,000
Member of Technical Staff, ML Compilers & Code Generation
Member of Technical Staff, ML Compilers & Code Generation

General Diffusion, Inc. • San Francisco (CA), Northern (KY)

Hybrid
USD 180,000 - 260,000
Member of Technical Staff, Distributed Systems & Fleet
Member of Technical Staff, Distributed Systems & Fleet

General Diffusion, Inc. • San Francisco (CA), Northern (KY)

Hybrid
USD 160,000 - 220,000
Research Engineer, Policy Evaluation
Research Engineer, Policy Evaluation

Bonfirevc • Palo Alto (CA)

On-site
USD 150,000 - 190,000
Research Engineer, RL Environments and Infrastructure
Research Engineer, RL Environments and Infrastructure

Hyphen Connect • San Francisco (CA)

Hybrid
USD 180,000 - 240,000
RL & Compute Environments Research Engineer
RL & Compute Environments Research Engineer

General Diffusion, Inc. • San Francisco (CA), Northern (KY)

Hybrid
USD 180,000 - 280,000
Staff Scientist, Action-Conditioned Compute World Models
Staff Scientist, Action-Conditioned Compute World Models

General Diffusion, Inc. • San Francisco (CA), Northern (KY)

Hybrid
USD 180,000 - 240,000