Lead AI Inference Systems Engineer

Hume AI

New York (NY)

On-site

USD 140,000 - 190,000

Full time

10 days ago
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Job summary

Hume AI in New York City is seeking a systems-oriented engineer to own the path from trained models to production inference. You will manage graph export, engine compilation, runtime integration, and serving contracts to ensure reliable, scalable inference.

You will work with research scientists, ML engineers, backend engineers, and the Data Plane team to deploy new models, optimize performance, and maintain observability and safety across the stack.

Qualifications

  • Significant professional experience building server-side, infrastructure, distributed, or systems software.
  • Strong Linux fundamentals and production troubleshooting.
  • Experience with a systems-oriented language such as Rust, Go, C, or C++.
  • Understanding of distributed systems concepts like load balancing, health checks, observability, and capacity management.
  • Practical understanding of neural-network execution, including graphs and tensor shapes.
  • Experience profiling and optimizing production systems.
  • Proficiency with Python for model export, validation, and integration workflows.
  • Strong ownership, independent judgment, and clear communication.

Responsibilities

  • Own the path from trained checkpoint to served requests, including graph export and engine integration.
  • Build and evolve the internal inference platform serving multiple models across products.
  • Design and operate serving infrastructure including routing, load balancing, health checks, and autoscaling.
  • Create versioned inference artifacts and tooling for validation, deployment, promotion, and rollback.
  • Implement verification gates to detect numerical, behavioral, and performance regressions before production.
  • Develop internal client libraries and standardized serving contracts.
  • Profile and optimize latency, throughput, memory usage, batching, and accelerator utilization.
  • Diagnose production issues across application, runtime, container, networking, driver, and hardware boundaries.
  • Improve observability, resilience, testability, and safety of the inference stack.
  • Write clear technical documentation for systems and APIs.

Skills

Systems software
Linux fundamentals
Rust/Go/C/C++
Distributed systems
Python for ML workflows
Ownership & communication
AI-assisted coding tools

Tools

CUDA
ONNX
TensorRT
Docker

Job description

Hume AI in New York City is seeking a systems-oriented engineer to own the path from trained models to production inference. You will manage graph export, engine compilation, runtime integration, and serving contracts to ensure reliable, scalable inference.

You will work with research scientists, ML engineers, backend engineers, and the Data Plane team to deploy new models, optimize performance, and maintain observability and safety across the stack.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior AI Inference Systems Engineer
Senior AI Inference Systems Engineer

Hume AI • New York (NY)

On-site
USD 180,000 - 260,000
Senior Systems Engineer, On-Prem AI Platform
Senior Systems Engineer, On-Prem AI Platform

Hume AI, Inc. • New York (NY), Northern (KY)

Hybrid
USD 150,000 - 230,000
Staff Software Engineer - Inference Backends
Staff Software Engineer - Inference Backends

Hume AI • New York (NY)

On-site
USD 140,000 - 190,000
Senior Software Engineer - Data Plane
Senior Software Engineer - Data Plane

Hume AI • New York (NY)

On-site
USD 180,000 - 260,000
Senior Software Engineer - Data Plane
Senior Software Engineer - Data Plane

Hume AI, Inc. • New York (NY), Northern (KY)

Hybrid
USD 150,000 - 230,000
Engineering Manager: Inference Infrastructure Leader
Engineering Manager: Inference Infrastructure Leader

EngineersOfAI • New York (NY), Northern (KY)

Hybrid
USD 230,000 - 360,000
Senior AI Inference Systems Engineer
Senior AI Inference Systems Engineer

Alex Loftus • Northern (KY), New York (NY)

Hybrid
USD 320,000 - 485,000
Equity donation matching
Vacation and parental leave
Flexible working hours
+1
Engineering Manager, ML Infrastructure & Fleet
Engineering Manager, ML Infrastructure & Fleet

Anthropic • New York (NY)

Hybrid
USD 405,000 - 625,000
Competitive compensation
Equity donation
Vacation and parental leave
+2
Senior AI Inference Engineer: High-Throughput LLMs
Senior AI Inference Engineer: High-Throughput LLMs

Crusoe Energy Systems • San Francisco (CA)

On-site
USD 180,000 - 240,000
Health benefits
401(k) match
Paid time off
+1
Senior Inference Engineer, AI Infrastructure & Production
Senior Inference Engineer, AI Infrastructure & Production

Hamilton Barnes • United States

On-site
USD 225,000 - 275,000
Full Benefits