Staff Engineer, AI Inference & Benchmarking

Liquid-Ai

Cambridge (MA)

Hybrid

USD 180,000 - 240,000

Full time

6 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Competitive base salary with equity
Health premiums paid (medical, dental,
401(k) matching up to 4%
Unlimited PTO and Refill Days

Job summary

Liquid AI in Boston, MA is seeking a Member Of Technical Staff, Infrastructure to drive the inference stack from research to production. You will design benchmark suites, port models across runtimes, and verify correctness end-to-end with external partners.

The role emphasizes C++ and Python in performance-sensitive contexts, with hybrid work and a strong focus on quantization, memory layout, and evaluation methodology to ensure reliable, scalable AI systems.

Qualifications

  • Hands-on with at least one inference framework like llama.cpp, ONNX Runtime, or MLX, going beyond basic usage into internals and modification.
  • Experience designing and building benchmarking pipelines, including methodology, validation, and reproducibility.
  • Strong C++ and Python in performance-sensitive contexts.
  • Solid understanding of inference fundamentals: quantization, decoding strategies, memory layout, and how they interact.

Responsibilities

  • Design and build benchmark suites that cover inference performance, model quality, and knowledge evaluation across different hardware targets.
  • Run external partner verifications: evaluate their solutions against our benchmarks, identify gaps, and clearly deliver findings.
  • Port models like LFM2 onto different runtimes and frameworks, and verify correctness end-to-end.
  • Maintain and extend the inference engine layer built on llama.cpp, ONNX, and MLX as new model architectures emerge from research.
  • Make benchmark results explainable and verifiable, so internal teams and partners can trust and reproduce them independently.

Skills

llama.cpp
ONNX Runtime
MLX
Benchmarking pipelines
C++
Python
Inference fundamentals

Job description

Liquid AI in Boston, MA is seeking a Member Of Technical Staff, Infrastructure to drive the inference stack from research to production. You will design benchmark suites, port models across runtimes, and verify correctness end-to-end with external partners.

The role emphasizes C++ and Python in performance-sensitive contexts, with hybrid work and a strong focus on quantization, memory layout, and evaluation methodology to ensure reliable, scalable AI systems.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Staff Engineer - AI Inference & Benchmarking
Staff Engineer - AI Inference & Benchmarking

Liquid-Ai • Boston (MA)

Hybrid
USD 140,000 - 210,000
Equity
Health insurance
401(k) match
+1
Member of Technical Staff - Inference Systems
Member of Technical Staff - Inference Systems

Liquid-Ai • Boston (MA)

Hybrid
USD 140,000 - 210,000
Equity
Health insurance
401(k) match
+1
Staff Engineer, AI Benchmarking & Inference Systems
Staff Engineer, AI Benchmarking & Inference Systems

S27a • San Francisco (CA), New York (NY)

Hybrid
USD 140,000 - 190,000
Generous PTO
Office stipend
Competitive healthcare (medical,Dental
+2
Staff Engineer, AI Inference & Benchmarking
Staff Engineer, AI Inference & Benchmarking

S27a • San Francisco (CA), New York (NY)

Hybrid
USD 150,000 - 210,000
Staff, AI Inference Benchmarking (Equity, Onsite SF)
Staff, AI Inference Benchmarking (Equity, Onsite SF)

Artificial Analysis, Inc. • San Francisco (CA)

On-site
USD 120,000 - 180,000
Equity
Member of Technical Staff - Inference Systems
Member of Technical Staff - Inference Systems

Liquid-Ai • Cambridge (MA)

Hybrid
USD 180,000 - 240,000
Competitive base salary with equity
Health premiums paid (medical, dental,
401(k) matching up to 4%
+1
Staff Engineer – High-Performance Model Inference
Staff Engineer – High-Performance Model Inference

Pantera Capital • Palo Alto (CA)

On-site
USD 180,000 - 440,000
Equity
Medical coverage
Vision coverage
+5
Inference Performance Engineer - Benchmark & Optimize
Inference Performance Engineer - Benchmark & Optimize

CoreWeave • Sunnyvale (CA)

On-site
USD 188,000 - 275,000
Medical, dental, and vision insurance
401(k) with employer match
Paid parental leave
+2
Inference Performance Engineer: Benchmark & Optimize
Inference Performance Engineer: Benchmark & Optimize

Coreweave • Bellevue (CA)

On-site
USD 188,000 - 275,000
Medical, dental, and vision insurance
Equity awards
401(k) with match
+1
Staff Engineer - Customer-Facing AI Inference Infra
Staff Engineer - Customer-Facing AI Inference Infra

Simplify • San Francisco (CA)

On-site
USD 200,000 - 300,000
Housing stipend
Uber/Waymo rides