Member of Technical Staff — Training Infrastructure

Human Intuition Inc.

New York (NY)

On-site

USD 150,000 - 210,000

Full time

14 days+
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Job summary

Human Intuition Inc. seeks to build the platform that turns research ideas into reproducible learning runs.

You will connect datasets, environments, rollout generation, trainers, checkpoints, and evaluations into a system researchers can understand and operate, shortening the path from a question to a trustworthy result. The role focuses on orchestration for fine-tuning and RL workloads, coordinating training and inference with reliable job state, and preserving experiment lineage across data,

Qualifications

  • Experience with ML infrastructure or distributed systems.
  • Strong Python and practical familiarity with training workloads.
  • Experience with workload orchestration, cloud infrastructure, or container platforms.
  • Ability to debug across application code, workers, networking, storage, and resource management.
  • Care for reproducibility and operational simplicity.

Responsibilities

  • Build orchestration for fine-tuning and reinforcement learning workloads across the compute resources the team uses.
  • Coordinate training, inference, and environment workers with clear job state and failure recovery.
  • Preserve experiment lineage across data, configuration, code, model versions, and evaluation results.
  • Improve checkpointing, artifact storage, job resumption, and resource utilization.
  • Provide useful logs, metrics, and debugging tools for learning and infrastructure failures.
  • Work with researchers to make new methods repeatable and with engineers to deliver validated models to serving systems.

Skills

Python
ML infrastructure
Distributed systems
Cloud platforms
Debugging across systems
Reproducibility

Tools

Kubernetes
Docker
Terraform
AWS

Job description

Building the autonomous company

Human Intuition is building the autonomous company. Businesses run on accumulated judgment: how to interpret a situation, choose an action, and learn from its consequences. Much of that knowledge lives in people, even when the decisions they make leave traces in software.

We are working to make that judgment learnable. A business has defined systems, tools, permissions, histories, and objectives. Those boundaries create an opportunity to build agents that learn from how work is done, act within clear constraints, and improve through feedback. Our ambition is to turn the knowledge inside institutions into software that compounds.

The role

Build the platform that turns research ideas into reproducible learning runs. You will connect datasets, environments, rollout generation, trainers, checkpoints, and evaluations into a system researchers can understand and operate. The aim is to shorten the path from a question to a trustworthy result.

What you’ll do
  • Build orchestration for fine-tuning and reinforcement learning workloads across the compute resources the team uses.

  • Coordinate training, inference, and environment workers with clear job state and failure recovery.

  • Preserve experiment lineage across data, configuration, code, model versions, and evaluation results.

  • Improve checkpointing, artifact storage, job resumption, and resource utilization.

  • Provide useful logs, metrics, and debugging tools for learning and infrastructure failures.

  • Work with researchers to make new methods repeatable and with engineers to deliver validated models to serving systems.

What you’ll bring
  • Experience with machine learning infrastructure or distributed systems used for compute-intensive work.

  • Strong Python and practical familiarity with training workloads and accelerator constraints.

  • Experience with workload orchestration, cloud infrastructure, or container platforms.

  • An ability to debug across application code, workers, networking, storage, and resource management.

  • Care for reproducibility, operational simplicity, and the experience of the people using the platform.

Useful experience

Distributed training, rollout systems, experiment tracking, model registries, GPU scheduling, or post-training infrastructure.

What success looks like

Researchers can launch, inspect, reproduce, and recover experiments with confidence, and successful results move into deployment with a clear record of how they were produced.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Member of Technical Staff — Inference
Member of Technical Staff — Inference

Human Intuition Inc. • New York (NY)

On-site
USD 140,000 - 195,000
Applied Research — RL & Agents
Applied Research — RL & Agents

Human Intuition Inc. • New York (NY)

On-site
USD 120,000 - 180,000
Member of Technical Staff — Sandbox Infrastructure
Member of Technical Staff — Sandbox Infrastructure

Human Intuition Inc. • New York (NY)

On-site
USD 110,000 - 170,000
Member of Technical Staff — Training Infrastructure
Member of Technical Staff — Training Infrastructure

Kindredventures • San Francisco (CA)

On-site
USD 180,000 - 240,000
Member of Technical Staff — Full Stack
Member of Technical Staff — Full Stack

Human Intuition Inc. • New York (NY)

On-site
USD 120,000 - 160,000
Member of Technical Staff, AI Compute & Data Infrastructure
Member of Technical Staff, AI Compute & Data Infrastructure

Vinci • Palo Alto (CA)

On-site
USD 180,000 - 240,000
Software Engineer - ML Infrastructure
Software Engineer - ML Infrastructure

Epsilon • San Francisco (CA), Northern (KY)

Hybrid
USD 180,000 - 280,000
Research Member of Technical Staff - Training Platform
Research Member of Technical Staff - Training Platform

Rhoda AI • Mountain View (CA)

On-site
USD 100,000 - 140,000
Member of Technical Staff
Member of Technical Staff

Harrison Clarke • San Francisco (CA)

On-site
USD 180,000 - 280,000
ML Research Engineer, Training
ML Research Engineer, Training

Weave Robotics • San Francisco (CA)

On-site
USD 180,000 - 250,000