Member of Technical Staff — Inference

Human Intuition Inc.

New York (NY)

On-site

USD 140,000 - 195,000

Full time

14 days+
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Job summary

Human Intuition Inc. is building the autonomous company and seeking an experienced engineer to ensure dependable, efficient model serving behind agents and training rollouts. You will work on serving, routing, and evaluating performance metrics that matter for task completion.

The role emphasizes reliability, scalability, and cost-aware design, with collaboration across research and infrastructure teams to support evaluation and post-training workloads.

Qualifications

  • Experience building and operating machine learning services or performance-sensitive distributed systems.
  • Strong Python and understanding of model serving, accelerator memory, and concurrency.
  • Hands-on experience with an inference engine or substantial production model-serving workloads.
  • Ability to investigate performance bottlenecks and distinguish measured improvements from assumptions.
  • Ownership of reliability, debugging, and clear operational documentation.

Responsibilities

  • Build inference services and model routing with explicit reliability, latency, and cost targets.
  • Profile representative agent workloads, including long contexts, tool calls, streaming, and concurrent requests.
  • Improve throughput and resource use through scheduling, batching, caching, and informed deployment choices.
  • Develop reproducible benchmarks that connect serving changes to task quality as well as speed and cost.
  • Implement versioned rollouts, observability, capacity planning, and practical failure recovery.

Skills

ML services
Distributed systems
Python
Performance optimization
Reliability ownership

Education

Bachelor's degree in CS/Math/ML

Tools

Inference engines
Profiling tools
Observability

Job description

Building the autonomous company

Human Intuition is building the autonomous company. Businesses run on accumulated judgment: how to interpret a situation, choose an action, and learn from its consequences. Much of that knowledge lives in people, even when the decisions they make leave traces in software.


We are working to make that judgment learnable. A business has defined systems, tools, permissions, histories, and objectives. Those boundaries create an opportunity to build agents that learn from how work is done, act within clear constraints, and improve through feedback. Our ambition is to turn the knowledge inside institutions into software that compounds.


The role

Make model execution dependable enough for business operations and efficient enough to improve continuously. You will build the serving and routing systems behind agents, evaluations, and training rollouts, and measure performance in terms that matter to complete tasks.


What you’ll do


  • Build inference services and model routing with explicit reliability, latency, and cost targets.


  • Profile representative agent workloads, including long contexts, tool calls, streaming, and concurrent requests.


  • Improve throughput and resource use through scheduling, batching, caching, and informed deployment choices.


  • Develop reproducible benchmarks that connect serving changes to task quality as well as speed and cost.


  • Implement versioned rollouts, observability, capacity planning, and practical failure recovery.


  • Partner with research and infrastructure engineers to support evaluation and post-training workloads.



What you’ll bring


  • Experience building and operating machine learning services or performance-sensitive distributed systems.


  • Strong Python and an understanding of model serving, accelerator memory, and concurrency.


  • Hands-on experience with an inference engine or substantial production model-serving workloads.


  • The ability to investigate performance bottlenecks and distinguish measured improvements from assumptions.


  • Ownership of reliability, debugging, and clear operational documentation.



Useful experience

GPU profiling, quantization, KV cache management, distributed serving, capacity scheduling, or contributions to inference software.


What success looks like

Agents and researchers have predictable access to models, and the team can explain and improve the cost and latency of completing a task.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Member of Technical Staff — Agent Systems
Member of Technical Staff — Agent Systems

Human Intuition Inc. • New York (NY)

On-site
USD 150,000 - 190,000
Applied Research — RL & Agents
Applied Research — RL & Agents

Human Intuition Inc. • New York (NY)

On-site
USD 120,000 - 180,000
INFERENCE ENGINEER
INFERENCE ENGINEER

MakerMaker.AI • San Francisco (CA)

On-site
USD 120,000 - 160,000
Member of Technical Staff — Training Infrastructure
Member of Technical Staff — Training Infrastructure

Human Intuition Inc. • New York (NY)

On-site
USD 150,000 - 210,000
Member of Technical Staff — Full Stack
Member of Technical Staff — Full Stack

Human Intuition Inc. • New York (NY)

On-site
USD 120,000 - 160,000
Applied Research — Evaluations & Data
Applied Research — Evaluations & Data

Human Intuition Inc. • New York (NY)

On-site
USD 120,000 - 180,000
Inference Performance Engineer
Inference Performance Engineer

Adaption Labs • San Francisco (CA)

On-site
USD 180,000 - 260,000
Flexible work
Adaption Passport
Lunch stipend
+1
Inference Engineer
Inference Engineer

Adaption Labs, Inc. • San Francisco (CA)

On-site
USD 180,000 - 280,000
Lunch stipend
Annual travel stipend
Flexible work options
+2
Member of Technical Staff, ML Engineer (Inference & Performance)
Member of Technical Staff, ML Engineer (Inference & Performance)

Bonfirevc • Palo Alto (CA)

On-site
USD 180,000 - 250,000
Applied Research — Forward Deployed
Applied Research — Forward Deployed

Human Intuition Inc. • New York (NY)

On-site
USD 150,000 - 190,000