Inference Systems Engineer: Benchmarking & Porting (Hybrid)

Liquid AI

Boston (MA)

Hybrid

USD 180,000 - 240,000

Full time

4 days ago
Be an early applicant
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Benefits offered by this job

Equity
Health insurance
401(k) matching
Unlimited PTO
Refill Days

Job summary

Liquid AI in Boston is seeking a Member Of Technical Staff, Infrastructure to join our Research & Engineering team. You will design benchmark suites and verify inference performance across hardware targets.

Port models like LFM2, extend the inference engine on llama.cpp, ONNX and MLX, and ensure correctness end-to-end while collaborating with research and product.

The role emphasizes hands-on work with C++ and Python in performance-sensitive contexts.

Qualifications

  • Hands-on experience with at least one inference framework like llama.cpp, ONNX Runtime, or MLX, going beyond basic usage into internals and modification.
  • Experience designing and building benchmarking pipelines, including methodology, validation, and reproducibility.
  • Strong C++ and Python in performance-sensitive contexts.
  • Solid understanding of inference fundamentals: quantization, decoding strategies, memory layout, and how they interact.

Responsibilities

  • Design and build benchmark suites that cover inference performance, model quality, and knowledge evaluation across different hardware targets.
  • Run external partner verifications: evaluate their solutions against our benchmarks, identify gaps, and clearly deliver findings.
  • Port models like LFM2 onto different runtimes and frameworks, and verify correctness end-to-end.
  • Maintain and extend the inference engine layer built on llama.cpp, ONNX, and MLX as new model architectures emerge from research.
  • Make benchmark results explainable and verifiable, so internal teams and partners can trust and reproduce them independently.

Skills

C++
Python
Inference frameworks
Benchmarking pipelines

Tools

llama.cpp
ONNX Runtime
MLX

Job description

Liquid AI in Boston is seeking a Member Of Technical Staff, Infrastructure to join our Research & Engineering team. You will design benchmark suites and verify inference performance across hardware targets.

Port models like LFM2, extend the inference engine on llama.cpp, ONNX and MLX, and ensure correctness end-to-end while collaborating with research and product.

The role emphasizes hands-on work with C++ and Python in performance-sensitive contexts.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Staff Engineer - AI Inference & Benchmarking
Staff Engineer - AI Inference & Benchmarking

Liquid-Ai • Boston (MA)

Hybrid
USD 140,000 - 210,000
Equity
Health insurance
401(k) match
+1
Staff Engineer, AI Inference & Benchmarking
Staff Engineer, AI Inference & Benchmarking

Liquid-Ai • Cambridge (MA)

Hybrid
USD 180,000 - 240,000
Competitive base salary with equity
Health premiums paid (medical, dental,
401(k) matching up to 4%
+1
Member of Technical Staff - Inference Systems
Member of Technical Staff - Inference Systems

Liquid AI • Boston (MA)

Hybrid
USD 180,000 - 240,000
Equity
Health insurance
401(k) matching
+2
Member of Technical Staff - Inference Systems
Member of Technical Staff - Inference Systems

Liquid-Ai • Boston (MA)

Hybrid
USD 140,000 - 210,000
Equity
Health insurance
401(k) match
+1
Member of Technical Staff - Inference Systems
Member of Technical Staff - Inference Systems

Liquid-Ai • Cambridge (MA)

Hybrid
USD 180,000 - 240,000
Competitive base salary with equity
Health premiums paid (medical, dental,
401(k) matching up to 4%
+1
Inference Systems Performance Engineer
Inference Systems Performance Engineer

Adaption Labs • San Francisco (CA), Northern (KY)

Hybrid
USD 180,000 - 260,000
Annual travel stipend
Lunch stipend
Well-Being benefits
Inference Systems Engineer — High-Performance ML Runtime
Inference Systems Engineer — High-Performance ML Runtime

The Consensus • San Jose (CA)

On-site
USD 180,000 - 240,000
Medical/dental/vision benefits
Housing subsidy
Relocation support
+2
Performance Engineer, Inference Engine - High-Performance AI
Performance Engineer, Inference Engine - High-Performance AI

EngineersOfAI • San Francisco (CA), Northern (KY)

Hybrid
USD 180,000 - 240,000
Technical Lead: Inference Benchmarking & ML Infra
Technical Lead: Inference Benchmarking & ML Infra

NVIDIA • United States

On-site
USD 224,000 - 357,000
Equity and benefits
Comprehensive benefits package
Competitive salaries
Senior LLM Infra Engineer — HPC, Benchmarking & DevEx
Senior LLM Infra Engineer — HPC, Benchmarking & DevEx

Baseten • United States

Remote
USD 180,000 - 300,000