Member of Technical Staff (Inference) - AI Infrastructure

Hamilton Barnes

United States

On-site

USD 225,000 - 275,000

Full time

6 days ago
Be an early applicant
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Benefits offered by this job

Full Benefits

Job summary

Hamilton Barnes is seeking an Inference Engineer to establish the technical foundation for inference, define architecture, and deploy production-ready systems. You will optimize model execution, memory management, and scheduling across machines, collaborating with Platform, Fleet, and Network Infrastructure teams.

The role covers kernels, compilers, inference engines, and serving APIs, balancing latency, throughput, model quality, reliability, and cost.

Qualifications

  • Production inference engines, GPU compute software, or distributed ML systems experience.
  • Strong C++ or Rust systems programming with Python proficiency.
  • Understanding transformer inference, including attention, batching, KV caches.
  • Experience with GPU programming, memory hierarchies, and performance analysis.
  • Distributed systems fundamentals: scheduling, concurrency, networking, and fault tolerance.
  • Experience delivering software from architecture to production deployment.
  • Ability to diagnose bottlenecks and implement measured improvements.
  • Willingness to own technical direction and communicate tradeoffs.

Responsibilities

  • Build and optimize the inference engine with model execution and memory management.
  • Write performance-critical kernels and runtime code across hardware and runtimes.
  • Design for high-performance hardware and profile on real devices.
  • Build distributed inference and its communication layer with model sharding and data movement.
  • Own inference scheduling, routing, and autoscaling to meet latency targets.
  • Ship a production inference service with loading, deployment, APIs, streaming, and backpressure.
  • Develop benchmarks to measure latency, throughput, and cost per token.
  • Set engineering direction and align with customer workloads and open-source projects.

Skills

Production inference engines
C++/Rust
Transformer inference
GPU programming
Distributed systems
Production deployment
Problem solving
Technical leadership

Job description

Ready to take the next step in your career?

Join a high-growth infrastructure company that makes Apple hardware available and performant at data centre scale for AI workloads, building cloud platforms that expose high-performance hardware as elastic compute through developer-friendly interfaces.

The organization is currently on the lookout for an Inference Engineer to establish the technical foundation for inference, defining the architecture, writing performance-critical code, and taking the system through deployment and production operation. The ideal candidate will work across kernels, compilers, inference engines, memory management, model parallelism, networking, scheduling, and serving APIs, collaborating closely with Platform, Fleet, and Network Infrastructure engineering to turn individual machines into a reliable inference service, owning the tradeoffs among latency, throughput, model quality, reliability, and cost.

Responsibilities:
  • Build and optimize the inference engine. Implement model execution, continuous batching, prefill and decode scheduling, KV-cache management, prefix caching, and memory allocation. Bring new model architectures into production and implement optimizations such as quantization, speculative decoding, and chunked prefill.
  • Write performance-critical kernels and runtime code. Implement and optimize operations such as matrix multiplication, attention, and mixture-of-experts execution. Work across Metal, MLX, and compiler and runtime internals to improve memory access, operator fusion, synchronization, and CPU/GPU execution. Validate numerical correctness alongside performance.
  • Design for Apple Silicon. Build around unified memory, memory bandwidth, compute capabilities, and operating-system behavior. Profile execution on real hardware, diagnose bottlenecks, and choose model representations and execution strategies based on measured results.
  • Build distributed inference and its communication layer. Implement model sharding and parallel execution across machines. Develop and optimize collective communication, data movement, and the overlap between computation and communication. Work with network engineers on transport performance, topology, and failure handling.
  • Own inference scheduling and routing. Build request queues, admission control, cache-aware routing, model placement, and autoscaling. Manage capacity and tenant fairness while meeting latency and throughput targets across changing traffic patterns and model sizes.
  • Ship a production inference service. Build model loading and deployment workflows, serving APIs, streaming responses, cancellation, and backpressure. Integrate the runtime with fleet and platform systems. Own observability, safe releases, incident diagnosis, and recovery for the inference stack.
  • Make performance reproducible. Build benchmarks and profiling tools that measure time to first token, inter-token latency, tail latency, throughput, memory use, power, and cost per token. Test realistic workloads, concurrency, and context lengths. Catch performance and model-quality regressions before release.
  • Set the engineering direction. Turn customer workloads into technical priorities, choose and contribute to open-source projects, and establish clear interfaces across the stack. Use coding agents to investigate, implement, and validate changes, backed by correctness checks and reproducible benchmarks.
Skills/Must Have:
  • A record of building and optimizing production inference engines, GPU compute software, or distributed ML systems, with substantial depth in at least one and hands-on work across multiple layers.
  • Strong systems programming skills in C++ or Rust, proficiency in Python, and experience working inside performance-critical libraries and runtimes.
  • A practical understanding of transformer inference, including attention, prefill and decode, batching, KV caches, quantization, and their compute and memory costs.
  • Experience with GPU programming and performance analysis, including memory hierarchies, parallel execution, synchronization, and numerical precision.
  • Strong distributed-systems fundamentals, including scheduling, concurrency, networking, failure handling, and resource management.
  • Experience taking software from architecture through production deployment, debugging, and ongoing operation.
  • The ability to turn an ambiguous performance problem into a measured bottleneck, an implementation, and a verified improvement.
  • The judgment and ownership to establish a new technical area, prioritize the work, and explain tradeoffs clearly to engineers and customers.
Benefits:
  • Full Benefits
Salary:
  • $250,000
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Member of Technical Staff, Inference
Member of Technical Staff, Inference

Mount Thor • San Francisco (CA)

On-site
USD 240,000 - 320,000
Inference Engineer
Inference Engineer

Acceler8 Talent • San Francisco (CA)

On-site
USD 180,000 - 220,000
Sr. AI Inference Platform Engineer
Sr. AI Inference Platform Engineer

Apple • Seattle (WA)

On-site
USD 175,000 - 309,000
Employee stock programs
Employee Stock Purchase Plan
Medical and dental coverage
+4
Sr. Machine Learning Engineer, Foundation Models Inference - Cloud OS & Inference
Sr. Machine Learning Engineer, Foundation Models Inference - Cloud OS & Inference

Apple Inc. • Santa Clara (CA), Northern (KY)

On-site
USD 185,000 - 325,000
Medical and dental coverage
Retirement benefits
Employee stock programs
Sr. Machine Learning Engineer, Foundation Models Inference - Cloud OS & Inference
Sr. Machine Learning Engineer, Foundation Models Inference - Cloud OS & Inference

Apple Inc. • Seattle (WA)

On-site
USD 185,000 - 325,000
Senior Software Engineer - Model Performance
Senior Software Engineer - Model Performance

inference.net • San Francisco (CA)

Hybrid
USD 220,000 - 320,000
Equity in a high-growth startup
Comprehensive benefits
Member of Technical Staff, Inference
Member of Technical Staff, Inference

Reactor • San Francisco (CA)

On-site
USD 180,000 - 240,000
Competitive SF salary
Early equity
Visa sponsorship
+2
INFERENCE OPTIMIZATION ENGINEER
INFERENCE OPTIMIZATION ENGINEER

Up Top • United States

Hybrid
USD 180,000 - 320,000
AI Inference Engineer
AI Inference Engineer

Premier Global Links • Palo Alto (CA)

On-site
USD 230,000 - 350,000
Senior Software Engineer - Model Performance
Senior Software Engineer - Model Performance

Inference • San Francisco (CA)

On-site
USD 220,000 - 320,000
Competitive compensation
Equity in a high-growth startup
Comprehensive benefits