Remote AI Inference Benchmark Engineer

Silicon Data

United States

Remote

USD 140,000 - 200,000

Full time

12 days ago
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Job summary

Silicon Data is seeking an AI Benchmark Engineer to build and run the inference performance side of SiliconMark. You will create automated, containerized benchmarking harnesses across engines and hardware, and own latency, throughput, and per-watt metrics.

This is a hands-on role requiring strong Python, experience with production inference engines (vLLM, TensorRT-LLM, SGLang, TGI), and deep knowledge of batching, KV cache, quantization, and parallelism. Remote work with ET overlap is offered.

Qualifications

  • 4+ years of engineering experience in LLM inference, model serving, or performance engineering.
  • Strong Python. You write automation and tooling that others can run and trust, not one-off scripts.
  • Direct experience with at least one production inference engine (vLLM, TensorRT-LLM, SGLang, TGI, or equivalent).
  • Working knowledge of what drives inference performance: batching/scheduling, KV cache, quantization, tensor/pipeline parallelism, memory bandwidth, prefill vs decode.

Responsibilities

  • Build and maintain the inference benchmarking harness across engines, models, and hardware targets.
  • Own inference metrics: latency, throughput, goodput, and cost per token.
  • Design workload profiles that reflect real usage patterns.
  • Capture and validate full configuration fingerprints for repeatable results.
  • Establish run repeatability and statistical rigor; manage outliers.

Skills

Python
Automation tooling
Inference performance
Benchmarking
Linux
Experimentation

Tools

vLLM
TensorRT-LLM
SGLang
TGI

Job description

Silicon Data is seeking an AI Benchmark Engineer to build and run the inference performance side of SiliconMark. You will create automated, containerized benchmarking harnesses across engines and hardware, and own latency, throughput, and per-watt metrics.

This is a hands-on role requiring strong Python, experience with production inference engines (vLLM, TensorRT-LLM, SGLang, TGI), and deep knowledge of batching, KV cache, quantization, and parallelism. Remote work with ET overlap is offered.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Staff Engineer, AI Inference & Benchmarking
Staff Engineer, AI Inference & Benchmarking

Liquid-Ai • Cambridge (MA)

Hybrid
USD 180,000 - 240,000
Competitive base salary with equity
Health premiums paid (medical, dental,
401(k) matching up to 4%
+1
Senior AI Inference Performance Engineer (Remote)
Senior AI Inference Performance Engineer (Remote)

DigitalOcean • Denver (CO)

On-site
USD 191,000 - 239,000
Staff Engineer - AI Inference & Benchmarking
Staff Engineer - AI Inference & Benchmarking

Liquid-Ai • Boston (MA)

Hybrid
USD 140,000 - 210,000
Equity
Health insurance
401(k) match
+1
AI Inference Benchmark Architect
AI Inference Benchmark Architect

Neurophos • Austin (TX), Sunnyvale (CA)

On-site
USD 150,000 - 210,000
Health insurance
Unlimited PTO
401k matching
+3
Staff Engineer - AI Workload Benchmarking
Staff Engineer - AI Workload Benchmarking

Prodapt ASIC services (Formerly Innovative Logic) • San Jose (CA)

On-site
USD 150,000 - 210,000
AI Benchmark Engineer, Inference Performance
AI Benchmark Engineer, Inference Performance

Silicon Data • United States

Remote
USD 140,000 - 200,000
Senior AI Inference Platform Engineer — Benchmarking
Senior AI Inference Platform Engineer — Benchmarking

Apple • Seattle (WA)

On-site
USD 175,000 - 309,000
Employee stock programs
Employee Stock Purchase Plan
Medical and dental coverage
+4
Staff Engineer - AI Benchmarking & Systems Performance
Staff Engineer - AI Benchmarking & Systems Performance

Prodapt ASIC services (Formerly Innovative Logic) • San Jose (CA)

On-site
USD 150,000 - 210,000
Hardware Benchmarking Engineer – AI Accelerators
Hardware Benchmarking Engineer – AI Accelerators

Artificial Analysis, Inc. • San Francisco (CA)

On-site
USD 150,000 - 210,000
Equity
Staff Engineer, AI Inference & Benchmarking
Staff Engineer, AI Inference & Benchmarking

S27a • San Francisco (CA), New York (NY)

Hybrid
USD 150,000 - 210,000