Senior LLM Inference Architect — Heterogeneous Hardware

d-Matrix inc.

Santa Clara (CA)

On-site

USD 130,000 - 170,000

Full time

14 days+
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Benefits offered by this job

Competitive compensation
Equity
Inclusive work environment

Job summary

d-Matrix inc. is hiring end-to-end inference engineers to innovate LLM solutions in Santa Clara, CA. You will work with novel hardware and software to optimize performance and deploy advanced inference systems.

Ideal candidates possess strong Python and C/C++ skills, along with a deep understanding of generative AI architectures and extensive experience in performance optimization of LLM frameworks. Join a small, senior team committed to cutting-edge AI development.

Qualifications

  • 10+ years of relevant engineering experience or equivalent.
  • Experience optimizing attention kernels, KV cache, batching strategies, and quantization (INT8/FP8/INT4).
  • Experience contributing to a major inference framework.

Responsibilities

  • Prototype emerging LLM inference use cases.
  • Build proof-of-concept systems.
  • Develop and tune custom kernels for optimal performance.

Skills

Python
C/C++
LLM inference optimization
CUDA/Triton
Performance profiling tools

Education

Bachelor's degree in Computer Science or Electrical Engineering
Master's or PhD preferred

Tools

vLLM
SGLang
TensorRT-LLM
ONNX Runtime

Job description

d-Matrix inc. is hiring end-to-end inference engineers to innovate LLM solutions in Santa Clara, CA. You will work with novel hardware and software to optimize performance and deploy advanced inference systems.

Ideal candidates possess strong Python and C/C++ skills, along with a deep understanding of generative AI architectures and extensive experience in performance optimization of LLM frameworks. Join a small, senior team committed to cutting-edge AI development.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior LLM Inference Engineer - End-to-End, Equity
Senior LLM Inference Engineer - End-to-End, Equity

Entrada Ventures • Santa Clara (CA)

On-site
USD 180,000 - 280,000
Competitive compensation
Equity opportunities
Benefits in Santa Clara, CA
Senior Staff LLM Inference Engineer
Senior Staff LLM Inference Engineer

Entrada Ventures • Santa Clara (CA)

On-site
USD 180,000 - 280,000
Competitive compensation
Equity opportunities
Benefits in Santa Clara, CA
LLM Inference Runtime Architect for AI Accelerator
LLM Inference Runtime Architect for AI Accelerator

United States Digital Space LLC • United States

Remote
USD 180,000 - 250,000
LLM AI Inference Performance Engineer
LLM AI Inference Performance Engineer

Intel • California (MO)

Hybrid
USD 171,000 - 315,000
LLM Inference Performance Engineer
LLM Inference Performance Engineer

Intel • Santa Clara (CA)

Hybrid
USD 171,000 - 315,000
Senior ML Researcher: LLM Inference & Algorithms Optimizer
Senior ML Researcher: LLM Inference & Algorithms Optimizer

Entrada Ventures • Santa Clara (CA)

Hybrid
USD 180,000 - 240,000
Senior ML Researcher - LLM Algorithmic Optimization
Senior ML Researcher - LLM Algorithmic Optimization

d-Matrix inc. • Santa Clara (CA)

Hybrid
USD 130,000 - 160,000
LLM Inference Frameworks & Optimizations Engineer
LLM Inference Frameworks & Optimizations Engineer

Together AI • San Francisco (CA)

On-site
USD 160,000 - 230,000
Equity
Health insurance
Benefits package
Senior ML Researcher: LLM Optimization (Hybrid, SC)
Senior ML Researcher: LLM Optimization (Hybrid, SC)

d-Matrix • United States

Hybrid
USD 180,000 - 230,000
Senior LLM Inference Engineer: Performance & Optimization
Senior LLM Inference Engineer: Performance & Optimization

Confidential • United States

On-site
USD 180,000 - 240,000