Senior LLM Inference Engineer - End-to-End, Equity

Entrada Ventures

Santa Clara (CA)

On-site

USD 180,000 - 280,000

Full time

14 days+
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Benefits offered by this job

Competitive compensation
Equity opportunities
Benefits in Santa Clara, CA

Job summary

d-Matrix Frontier Group in Santa Clara, CA is seeking end-to-end inference engineers to drive novel ideas to deployed, optimized systems across the inference stack. You will work from kernel-level optimization to distributed orchestration and high-level serving APIs, shaping the future of AI silicon usage.

Ideal candidates have deep experience with LLM inference, open-source frameworks, and heterogeneous hardware deployments, delivering POCs to customers and contributing to open-source projects.

Qualifications

  • Bachelors with 10+ years of relevant engineering experience or equivalent.
  • Masters/PhD with 6+ years of relevant industry experience preferred.
  • Strong Python and C/C++ proficiency.
  • Hands-on experience optimizing LLM inference and related kernels.
  • Experience with major inference frameworks at contributor level.
  • Familiarity with GPU kernel programming and profiling tools.

Responsibilities

  • Identify and prototype emerging LLM inference use cases for heterogeneous hardware.
  • Build POCs demonstrating D-Matrix capabilities to customers and partners.
  • Develop and tune custom kernels and operator-level optimizations for throughput and latency.
  • Drive quantization, sparsity, and batching strategies for the computational model.
  • Build and maintain inference runtimes, serving frameworks, and evaluation tooling.
  • Contribute to distributed inference systems and work with hardware teams.

Skills

Python
C/C++
LLM inference
CUDA/Triton
Performance profiling

Education

Bachelor's degree in CS/EE
Master's or PhD in CS/EE

Tools

vLLM
SGLang
TensorRT-LLM
ONNX Runtime

Job description

d-Matrix Frontier Group in Santa Clara, CA is seeking end-to-end inference engineers to drive novel ideas to deployed, optimized systems across the inference stack. You will work from kernel-level optimization to distributed orchestration and high-level serving APIs, shaping the future of AI silicon usage.

Ideal candidates have deep experience with LLM inference, open-source frameworks, and heterogeneous hardware deployments, delivering POCs to customers and contributing to open-source projects.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior LLM Inference Architect — Heterogeneous Hardware
Senior LLM Inference Architect — Heterogeneous Hardware

d-Matrix inc. • Santa Clara (CA)

On-site
USD 130,000 - 170,000
Competitive compensation
Equity
Inclusive work environment
LLM Inference Performance Engineer
LLM Inference Performance Engineer

Intel • Santa Clara (CA)

Hybrid
USD 171,000 - 315,000
Senior Staff LLM Inference Engineer
Senior Staff LLM Inference Engineer

Entrada Ventures • Santa Clara (CA)

On-site
USD 180,000 - 280,000
Competitive compensation
Equity opportunities
Benefits in Santa Clara, CA
LLM Inference Engineer — Distributed Systems
LLM Inference Engineer — Distributed Systems

OpenTalent • San Francisco (CA)

On-site
USD 180,000 - 260,000
Distributed LLM Inference Engineer - Scale HighThroughput AI
Distributed LLM Inference Engineer - Scale HighThroughput AI

Cerebras • Palo Alto (CA)

On-site
USD 120,000 - 160,000
Stock Options
Healthcare plans with 99% premium coverage
401k Retirement Plan
+6
LLM AI Inference Performance Engineer
LLM AI Inference Performance Engineer

Intel • California (MO)

Hybrid
USD 171,000 - 315,000
LLM Inference Optimization Engineer — Frontier Performance
LLM Inference Optimization Engineer — Frontier Performance

NLP PEOPLE • Sonoma (CA)

On-site
USD 120,000 - 160,000
Senior ML Researcher: LLM Inference & Algorithms Optimizer
Senior ML Researcher: LLM Inference & Algorithms Optimizer

Entrada Ventures • Santa Clara (CA)

Hybrid
USD 180,000 - 240,000
Senior ML Researcher: LLM Optimization (Hybrid, SC)
Senior ML Researcher: LLM Optimization (Hybrid, SC)

d-Matrix • United States

Hybrid
USD 180,000 - 230,000
LLM Inference Runtime Architect for AI Accelerator
LLM Inference Runtime Architect for AI Accelerator

United States Digital Space LLC • United States

Remote
USD 180,000 - 250,000