Senior Inference Engineer, AI Infrastructure & Production

Hamilton Barnes

United States

On-site

USD 225,000 - 275,000

Full time

5 days ago
Be an early applicant
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Benefits offered by this job

Full Benefits

Job summary

Hamilton Barnes is seeking an Inference Engineer to establish the technical foundation for inference, define architecture, and deploy production-ready systems. You will optimize model execution, memory management, and scheduling across machines, collaborating with Platform, Fleet, and Network Infrastructure teams.

The role covers kernels, compilers, inference engines, and serving APIs, balancing latency, throughput, model quality, reliability, and cost.

Qualifications

  • Production inference engines, GPU compute software, or distributed ML systems experience.
  • Strong C++ or Rust systems programming with Python proficiency.
  • Understanding transformer inference, including attention, batching, KV caches.
  • Experience with GPU programming, memory hierarchies, and performance analysis.
  • Distributed systems fundamentals: scheduling, concurrency, networking, and fault tolerance.
  • Experience delivering software from architecture to production deployment.
  • Ability to diagnose bottlenecks and implement measured improvements.
  • Willingness to own technical direction and communicate tradeoffs.

Responsibilities

  • Build and optimize the inference engine with model execution and memory management.
  • Write performance-critical kernels and runtime code across hardware and runtimes.
  • Design for high-performance hardware and profile on real devices.
  • Build distributed inference and its communication layer with model sharding and data movement.
  • Own inference scheduling, routing, and autoscaling to meet latency targets.
  • Ship a production inference service with loading, deployment, APIs, streaming, and backpressure.
  • Develop benchmarks to measure latency, throughput, and cost per token.
  • Set engineering direction and align with customer workloads and open-source projects.

Skills

Production inference engines
C++/Rust
Transformer inference
GPU programming
Distributed systems
Production deployment
Problem solving
Technical leadership

Job description

Hamilton Barnes is seeking an Inference Engineer to establish the technical foundation for inference, define architecture, and deploy production-ready systems. You will optimize model execution, memory management, and scheduling across machines, collaborating with Platform, Fleet, and Network Infrastructure teams.

The role covers kernels, compilers, inference engines, and serving APIs, balancing latency, throughput, model quality, reliability, and cost.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Staff Engineer – High-Performance Model Inference
Staff Engineer – High-Performance Model Inference

Pantera Capital • Palo Alto (CA)

On-site
USD 180,000 - 440,000
Equity
Medical coverage
Vision coverage
+5
Senior Platform Engineer, Inference & GPU Compute Infra
Senior Platform Engineer, Inference & GPU Compute Infra

Together • San Francisco (CA)

On-site
USD 240,000 - 280,000
Startup equity
Health insurance
INFERENCE ENGINEER
INFERENCE ENGINEER

MakerMaker.AI • San Francisco (CA)

On-site
USD 120,000 - 160,000
Engineering Manager: Inference Infrastructure Leader
Engineering Manager: Inference Infrastructure Leader

EngineersOfAI • New York (NY), Northern (KY)

Hybrid
USD 230,000 - 360,000
Inference Systems Engineer — High-Performance AI Serving
Inference Systems Engineer — High-Performance AI Serving

River AI • Palo Alto (CA)

On-site
USD 200,000 - 420,000
Health insurance
Dental insurance
Vision insurance
+3
Staff Engineer, AI Inference & Benchmarking
Staff Engineer, AI Inference & Benchmarking

Liquid-Ai • Cambridge (MA)

Hybrid
USD 180,000 - 240,000
Competitive base salary with equity
Health premiums paid (medical, dental,
401(k) matching up to 4%
+1
Junior AI Infrastructure Engineer - Build Scalable Inference
Junior AI Infrastructure Engineer - Build Scalable Inference

Neural Solutions • Columbia (MD)

On-site
USD 138,000 - 163,000
Member of Technical Staff (Inference) - AI Infrastructure
Member of Technical Staff (Inference) - AI Infrastructure

Hamilton Barnes • United States

On-site
USD 225,000 - 275,000
Full Benefits
Senior Inference Systems Engineer for High-Performance AI
Senior Inference Systems Engineer for High-Performance AI

Causal Labs • San Francisco (CA)

On-site
USD 150,000 - 210,000
Senior AI Inference Platform Engineer
Senior AI Inference Platform Engineer

Cloudflare • Austin (TX)

Hybrid
USD 180,000 - 240,000