Turn this role into an interview — a resume and cover letter built around what this employer wants.
Human Intuition Inc. is building the autonomous company and seeking an experienced engineer to ensure dependable, efficient model serving behind agents and training rollouts. You will work on serving, routing, and evaluating performance metrics that matter for task completion.
The role emphasizes reliability, scalability, and cost-aware design, with collaboration across research and infrastructure teams to support evaluation and post-training workloads.
Human Intuition is building the autonomous company. Businesses run on accumulated judgment: how to interpret a situation, choose an action, and learn from its consequences. Much of that knowledge lives in people, even when the decisions they make leave traces in software.
We are working to make that judgment learnable. A business has defined systems, tools, permissions, histories, and objectives. Those boundaries create an opportunity to build agents that learn from how work is done, act within clear constraints, and improve through feedback. Our ambition is to turn the knowledge inside institutions into software that compounds.
Make model execution dependable enough for business operations and efficient enough to improve continuously. You will build the serving and routing systems behind agents, evaluations, and training rollouts, and measure performance in terms that matter to complete tasks.
Build inference services and model routing with explicit reliability, latency, and cost targets.
Profile representative agent workloads, including long contexts, tool calls, streaming, and concurrent requests.
Improve throughput and resource use through scheduling, batching, caching, and informed deployment choices.
Develop reproducible benchmarks that connect serving changes to task quality as well as speed and cost.
Implement versioned rollouts, observability, capacity planning, and practical failure recovery.
Partner with research and infrastructure engineers to support evaluation and post-training workloads.
Experience building and operating machine learning services or performance-sensitive distributed systems.
Strong Python and an understanding of model serving, accelerator memory, and concurrency.
Hands-on experience with an inference engine or substantial production model-serving workloads.
The ability to investigate performance bottlenecks and distinguish measured improvements from assumptions.
Ownership of reliability, debugging, and clear operational documentation.
GPU profiling, quantization, KV cache management, distributed serving, capacity scheduling, or contributions to inference software.
Agents and researchers have predictable access to models, and the team can explain and improve the cost and latency of completing a task.