ML Inference & Performance Engineer

Bonfirevc

Palo Alto (CA)

On-site

USD 180,000 - 250,000

Full time

9 days ago
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Job summary

Orbifold AI in Palo Alto is seeking a Member of Technical Staff, ML Engineer (Inference & Performance) to own the end-to-end serving layer for production models on Ray Serve and PyTorch, across a heterogeneous GPU fleet. This role focuses on throughput, latency, memory fit, and cost efficiency.

You will push quantization, memory layout, and kernel-level optimizations, with opportunities to develop custom CUDA/Triton code, benchmarks, and production-grade observability for scalable, reliable

Qualifications

  • 3+ years building production ML systems with Python and PyTorch.
  • Experience owning inference performance for a real workload with measurable improvements.
  • Understanding of GPU execution: memory hierarchy, bandwidth, and kernel launch overhead.
  • Experience operating a distributed serving framework such as Ray or Kubernetes in production.
  • Comfortable being measured on throughput, latency, and cost.
  • Judgment about when to optimize and when to leave something alone.

Responsibilities

  • Own the serving end-to-end for production models on Ray Serve and PyTorch across a heterogeneous GPU fleet.
  • Design batching strategies, scheduling, concurrency, queueing, and related performance optimizations.
  • Implement quantization, memory layout, and cache management while validating output quality.
  • Develop compiler paths or custom CUDA/Triton kernels where beneficial.
  • Serve high-volume video and multimodal inference workloads with fast rollout cycles.
  • Build a benchmarking harness for reproducible performance claims and to catch regressions.
  • Operate in production with autoscaling, fault tolerance, and observability.
  • Onboard partner models onto our infrastructure and ensure parity with local results.

Skills

Python
PyTorch
Production ML systems
Ray
Kubernetes
GPU execution
Throughput optimization

Tools

Ray Serve
TensorRT-LLM
Torch.compile

Job description

Orbifold AI in Palo Alto is seeking a Member of Technical Staff, ML Engineer (Inference & Performance) to own the end-to-end serving layer for production models on Ray Serve and PyTorch, across a heterogeneous GPU fleet. This role focuses on throughput, latency, memory fit, and cost efficiency.

You will push quantization, memory layout, and kernel-level optimizations, with opportunities to develop custom CUDA/Triton code, benchmarks, and production-grade observability for scalable, reliable

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Staff ML Engineer — Production Inference & Serving
Staff ML Engineer — Production Inference & Serving

Orbifold AI, Inc. • Palo Alto (CA)

On-site
USD 180,000 - 240,000
Member of Technical Staff, ML Engineer (Inference & Performance)
Member of Technical Staff, ML Engineer (Inference & Performance)

Bonfirevc • Palo Alto (CA)

On-site
USD 180,000 - 250,000
ML Inference Engineer - High-Performance AI Systems
ML Inference Engineer - High-Performance AI Systems

Together Computer Inc • San Francisco (CA)

On-site
USD 200,000 - 300,000
Startup equity
Health insurance
Competitive compensation
ML Inference Performance Engineer — Optimize Cost & Latency
ML Inference Performance Engineer — Optimize Cost & Latency

Adaption Labs • San Francisco (CA)

On-site
USD 180,000 - 260,000
Flexible work
Adaption Passport
Lunch stipend
+1
Machine Learning Engineer (Inference)
Machine Learning Engineer (Inference)

Acceler8 Talent • San Francisco (CA), Northern (KY)

Hybrid
USD 180,000 - 240,000
Senior ML Engineer: AI Inference & Performance Optimizer
Senior ML Engineer: AI Inference & Performance Optimizer

Nebius • Palo Alto (CA)

Hybrid
USD 195,200 - 262,200
Health insurance
401(k) plan
Parental leave
+2
ML Inference Engineer San Francisco · Engineering · Full Time →
ML Inference Engineer San Francisco · Engineering · Full Time →

Reactor • San Francisco (CA)

On-site
USD 120,000 - 160,000
Visa sponsorship
Relocation support
Generous health, dental, and vision coverage
AI Inference Engineer
AI Inference Engineer

Acceler8 Talent • San Francisco (CA)

On-site
USD 150,000 - 230,000
High-Performance ML Inference Engineer
High-Performance ML Inference Engineer

Reactor • San Francisco (CA)

On-site
USD 180,000 - 240,000
Competitive SF salary
Early equity
Visa sponsorship
+2
Member of Technical Staff, Inference & Serving
Member of Technical Staff, Inference & Serving

Inception • San Francisco (CA)

On-site
USD 180,000 - 240,000