An application made for this job — a tailored resume and cover letter that speak straight to the posting.
Human Intuition Inc. seeks to build the platform that turns research ideas into reproducible learning runs.
You will connect datasets, environments, rollout generation, trainers, checkpoints, and evaluations into a system researchers can understand and operate, shortening the path from a question to a trustworthy result. The role focuses on orchestration for fine-tuning and RL workloads, coordinating training and inference with reliable job state, and preserving experiment lineage across data,
Human Intuition is building the autonomous company. Businesses run on accumulated judgment: how to interpret a situation, choose an action, and learn from its consequences. Much of that knowledge lives in people, even when the decisions they make leave traces in software.
We are working to make that judgment learnable. A business has defined systems, tools, permissions, histories, and objectives. Those boundaries create an opportunity to build agents that learn from how work is done, act within clear constraints, and improve through feedback. Our ambition is to turn the knowledge inside institutions into software that compounds.
Build the platform that turns research ideas into reproducible learning runs. You will connect datasets, environments, rollout generation, trainers, checkpoints, and evaluations into a system researchers can understand and operate. The aim is to shorten the path from a question to a trustworthy result.
Build orchestration for fine-tuning and reinforcement learning workloads across the compute resources the team uses.
Coordinate training, inference, and environment workers with clear job state and failure recovery.
Preserve experiment lineage across data, configuration, code, model versions, and evaluation results.
Improve checkpointing, artifact storage, job resumption, and resource utilization.
Provide useful logs, metrics, and debugging tools for learning and infrastructure failures.
Work with researchers to make new methods repeatable and with engineers to deliver validated models to serving systems.
Experience with machine learning infrastructure or distributed systems used for compute-intensive work.
Strong Python and practical familiarity with training workloads and accelerator constraints.
Experience with workload orchestration, cloud infrastructure, or container platforms.
An ability to debug across application code, workers, networking, storage, and resource management.
Care for reproducibility, operational simplicity, and the experience of the people using the platform.
Distributed training, rollout systems, experiment tracking, model registries, GPU scheduling, or post-training infrastructure.
Researchers can launch, inspect, reproduce, and recover experiments with confidence, and successful results move into deployment with a clear record of how they were produced.