Staff ML Systems Engineer – Reliable Inference & Serving

Human Intuition Inc.

New York (NY)

On-site

USD 140,000 - 195,000

Full time

14 days+
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Job summary

Human Intuition Inc. is building the autonomous company and seeking an experienced engineer to ensure dependable, efficient model serving behind agents and training rollouts. You will work on serving, routing, and evaluating performance metrics that matter for task completion.

The role emphasizes reliability, scalability, and cost-aware design, with collaboration across research and infrastructure teams to support evaluation and post-training workloads.

Qualifications

  • Experience building and operating machine learning services or performance-sensitive distributed systems.
  • Strong Python and understanding of model serving, accelerator memory, and concurrency.
  • Hands-on experience with an inference engine or substantial production model-serving workloads.
  • Ability to investigate performance bottlenecks and distinguish measured improvements from assumptions.
  • Ownership of reliability, debugging, and clear operational documentation.

Responsibilities

  • Build inference services and model routing with explicit reliability, latency, and cost targets.
  • Profile representative agent workloads, including long contexts, tool calls, streaming, and concurrent requests.
  • Improve throughput and resource use through scheduling, batching, caching, and informed deployment choices.
  • Develop reproducible benchmarks that connect serving changes to task quality as well as speed and cost.
  • Implement versioned rollouts, observability, capacity planning, and practical failure recovery.

Skills

ML services
Distributed systems
Python
Performance optimization
Reliability ownership

Education

Bachelor's degree in CS/Math/ML

Tools

Inference engines
Profiling tools
Observability

Job description

Human Intuition Inc. is building the autonomous company and seeking an experienced engineer to ensure dependable, efficient model serving behind agents and training rollouts. You will work on serving, routing, and evaluating performance metrics that matter for task completion.

The role emphasizes reliability, scalability, and cost-aware design, with collaboration across research and infrastructure teams to support evaluation and post-training workloads.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Inference Systems Engineer — Scalable, Reliable ML Serving
Inference Systems Engineer — Scalable, Reliable ML Serving

Speedrun Talent Network • San Francisco (CA)

On-site
USD 300,000 - 400,000
Health, dental, and vision benefits
Unlimited PTO
Paid parental leave
+1
Member of Technical Staff — Inference
Member of Technical Staff — Inference

Human Intuition Inc. • New York (NY)

On-site
USD 140,000 - 195,000
Senior Inference & RL Systems Engineer (Scalable ML Infra)
Senior Inference & RL Systems Engineer (Scalable ML Infra)

Magic AI, Inc • San Francisco (CA)

On-site
USD 300,000 - 550,000
Equity compensation
401(k) matching
Health, dental and vision insurance
+4
Production ML Inference Engineer — Scale & Reliability
Production ML Inference Engineer — Scale & Reliability

AI Chopping Block, Inc. • San Francisco (CA)

On-site
USD 300,000 - 400,000
Health, dental, and vision benefits
Unlimited PTO
Paid parental leave
+1
Staff Engineer, Inference & RL Systems — Scale ML
Staff Engineer, Inference & RL Systems — Scale ML

magic.dev • San Francisco (CA)

On-site
USD 275,000 - 550,000
Equity
401(k) matching
Health insurance
+4
Senior ML Model Serving Engineer - Remote
Senior ML Model Serving Engineer - Remote

Bright Vision Technologies • United States

Remote
USD 74,000 - 98,000
Real-Time ML Inference Engineer for Scalable Serving
Real-Time ML Inference Engineer for Scalable Serving

Yobi • New York (NY)

Hybrid
USD 100,000 - 150,000
Competitive Base Salary
Meaningful equity
Annual performance bonus
+3
Realtime ML Inference Engineer — Scalable Serving
Realtime ML Inference Engineer — Scalable Serving

Yobi AI • New York (NY)

Remote
USD 120,000 - 150,000
Competitive base salary
Meaningful equity
Annual bonus
+3
ML Inference Systems Engineer — High-Performance Serving
ML Inference Systems Engineer — High-Performance Serving

Annapurna Labs (U.S.) Inc. • Seattle (WA)

On-site
USD 144,000 - 194,000
RSUs
401(k) matching
Paid time off
Staff ML Inference Systems Engineer, Real-Time Safety
Staff ML Inference Systems Engineer, Real-Time Safety

EngineersOfAI • San Francisco (CA), Northern (KY)

Hybrid
USD 272,000 - 368,000