Member of Technical Staff - Inference Systems

Liquid-Ai

Cambridge (MA)

Hybrid

USD 180,000 - 240,000

Full time

14 days+
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Benefits offered by this job

Competitive base salary with equity
Health premiums paid (medical, dental,
401(k) matching up to 4%
Unlimited PTO and Refill Days

Job summary

Liquid AI in Boston, MA is seeking a Member Of Technical Staff, Infrastructure to drive the inference stack from research to production. You will design benchmark suites, port models across runtimes, and verify correctness end-to-end with external partners.

The role emphasizes C++ and Python in performance-sensitive contexts, with hybrid work and a strong focus on quantization, memory layout, and evaluation methodology to ensure reliable, scalable AI systems.

Qualifications

  • Hands-on with at least one inference framework like llama.cpp, ONNX Runtime, or MLX, going beyond basic usage into internals and modification.
  • Experience designing and building benchmarking pipelines, including methodology, validation, and reproducibility.
  • Strong C++ and Python in performance-sensitive contexts.
  • Solid understanding of inference fundamentals: quantization, decoding strategies, memory layout, and how they interact.

Responsibilities

  • Design and build benchmark suites that cover inference performance, model quality, and knowledge evaluation across different hardware targets.
  • Run external partner verifications: evaluate their solutions against our benchmarks, identify gaps, and clearly deliver findings.
  • Port models like LFM2 onto different runtimes and frameworks, and verify correctness end-to-end.
  • Maintain and extend the inference engine layer built on llama.cpp, ONNX, and MLX as new model architectures emerge from research.
  • Make benchmark results explainable and verifiable, so internal teams and partners can trust and reproduce them independently.

Skills

llama.cpp
ONNX Runtime
MLX
Benchmarking pipelines
C++
Python
Inference fundamentals

Job description

Liquid AI Job Description

Role: Member Of Technical Staff, Infrastructure

Department: Research & Engineering

Location: Boston

Location Type: Hybrid

Employment Type: Full-time

About Liquid AI

Spun out of MIT CSAIL, we build general-purpose AI systems that run efficiently across deployment targets, from data center accelerators to on-device hardware, ensuring low latency, minimal memory usage, privacy, and reliability. We partner with enterprises across consumer electronics, automotive, life sciences, and financial services. We are scaling rapidly and need exceptional people to help us get there.

The Opportunity

Our inference stack is central to everything we ship. You'll be a core part of the team responsible for the engine layer that runs our models in production and in partner environments, and for the benchmarking infrastructure we use to evaluate our own work and verify what partners bring to us. Day to day, that means working closely with research and product, but also directly with external engineering teams.

What We're Looking For

We need someone who:

  • Can pick up unfamiliar tools quickly and knows how to assess whether they're worth using.

  • Designs AI benchmarks and holds methodology to a high standard.

  • Cares about inference details, understands the tradeoffs, and checks what changed across the board before calling something done.

  • Doesn’t consider a model port finished until you can prove the outputs are correct.

The Work
  • Design and build benchmark suites that cover inference performance, model quality, and knowledge evaluation across different hardware targets.

  • Run external partner verifications: evaluate their solutions against our benchmarks, identify gaps, and clearly deliver findings.

  • Port models like LFM2 onto different runtimes and frameworks, and verify correctness end-to-end.

  • Maintain and extend the inference engine layer built on llama.cpp, ONNX, and MLX as new model architectures emerge from research.

  • Make benchmark results explainable and verifiable, so internal teams and partners can trust and reproduce them independently.

Desired Experience

Must-have:

  • Hands-on experience with at least one inference framework like llama.cpp, ONNX Runtime, or MLX, going beyond basic usage into internals and modification.

  • Experience designing and building benchmarking pipelines, including methodology, validation, and reproducibility.

  • Strong C++ and Python in performance-sensitive contexts.

  • Solid understanding of inference fundamentals: quantization, decoding strategies, memory layout, and how they interact.

Nice-to-have:

  • Experience porting models across runtimes and verifying numerical correctness.

  • Prior work with external partners or clients in a technical validation or evaluation capacity.

  • Familiarity with edge inference targets and the constraints that come with them.

What Success Looks Like (Year One)
  • You've ported LFM2 onto multiple runtimes and platforms, you know the model inside out, and new ports take you a fraction of the time they did at the start.

  • You've run multiple partner verifications end-to-end and built enough context to spot weak evaluations quickly and push back with evidence.

  • The benchmark suite covers inference performance and model quality across the platforms we care about, and both internal teams and partners are using it as a reference.

What We Offer
  • Compensation: Competitive base salary with equity in a unicorn-stage company

  • Health: We pay 100% of medical, dental, and vision premiums for employees and dependents

  • Financial: 401(k) matching up to 4% of base pay

  • Time Off: Unlimited PTO plus company-wide Refill Days throughout the year

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Member of Technical Staff - Inference Systems
Member of Technical Staff - Inference Systems

Liquid-Ai • Boston (MA)

Hybrid
USD 140,000 - 210,000
Equity
Health insurance
401(k) match
+1
Solutions Architect
Solutions Architect

Liquid AI • San Francisco (CA)

On-site
USD 120,000 - 160,000
Competitive base salary
Equity in a unicorn-stage company
100% medical, dental, and vision premiums
+2
Member of Technical Staff - Post Training, Applied (Text)
Member of Technical Staff - Post Training, Applied (Text)

Liquid AI • San Francisco (CA)

On-site
USD 100,000 - 140,000
Competitive base salary with equity
100% paid health, dental, and vision premiums
401(k) matching up to 4%
+1
Member of Technical Staff (AI Inference Engineer)
Member of Technical Staff (AI Inference Engineer)

Kindredventures • Palo Alto (CA)

On-site
USD 190,000 - 250,000
Comprehensive health insurance
Dental insurance
Vision insurance
+1
Software Engineer, Inference Runtime
Software Engineer, Inference Runtime

Lm-Studio • New York (NY)

Hybrid
USD 140,000 - 210,000
Competitive salary and equity grants
Excellent medical, vision, dental care
Catered team lunch / expensed dinners
+3
Distributed Systems Engineer, Data & Inference Platform
Distributed Systems Engineer, Data & Inference Platform

OpenTalent • San Francisco (CA)

On-site
USD 150,000 - 230,000
Flexible work
Adaption Passport
Lunch Stipend
+1
Member of Technical Staff — Model Optimization and Inference (New Grad)
Member of Technical Staff — Model Optimization and Inference (New Grad)

Nuance Labs • Seattle (WA)

On-site
USD 200,000 - 300,000
Health Savings Account with $2,000 annual contributions
15 days of PTO plus public holidays
Lunch, drinks, and snacks provided daily
Member of Technical Staff - Embedded ML Engineer (Audio/Omni)
Member of Technical Staff - Embedded ML Engineer (Audio/Omni)

Doist • San Francisco (CA)

On-site
USD 140,000 - 220,000
Health insurance
401(k) matching
Unlimited PTO
+1
Member of Technical Staff (Inference)
Member of Technical Staff (Inference)

Artificial Analysis, Inc. • San Francisco (CA)

On-site
USD 120,000 - 180,000
Equity
Member of Technical Staff - Embedded ML Engineer (Audio/Omni)
Member of Technical Staff - Embedded ML Engineer (Audio/Omni)

Liquid AI • United States

On-site
USD 120,000 - 180,000
Health premiums covered
401(k) matching
Unlimited PTO