Principal LLM Inference Engineer

Entrada Ventures

Santa Clara (CA)

On-site

USD 120,000 - 160,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Competitive compensation
Equity
Benefits

Job summary

Entrada Ventures in Santa Clara is seeking end-to-end inference engineers to work on innovative AI architecture. Candidates will enjoy ownership from research to deployment, tackling optimization challenges unique to D-Matrix's groundbreaking hardware.

Excellent communication, experience with AI frameworks, and strong proficiency in Python and C/C++ are essential. Other benefits include competitive compensation and opportunity for influence within a senior team.

Qualifications

  • Bachelor's degree in a related field with 10+ years of engineering experience.
  • Master's or PhD preferred with 6+ years relevant industry experience.
  • Strong proficiency in Python and C/C++.

Responsibilities

  • Identify and prototype emerging LLM inference use cases for heterogeneous hardware.
  • Build proof-of-concept systems demonstrating D-Matrix capabilities.
  • Contribute to distributed inference systems and partner with product development.

Skills

Python
C/C++
LLM inference optimization
CUDA
Performance profiling tools

Education

Bachelor's in Computer Science or Electrical Engineering
Master's or PhD in related field

Tools

vLLM
SGLang
TensorRT-LLM
ONNX Runtime

Job description

Role Overview

We are hiring end-to-end inference engineers who are comfortable going from a novel research idea to a deployed, optimized system. You will work at every layer of the inference stack — from kernel-level optimization to distributed orchestration to high-level serving APIs.

This role could be a great match for you if you:
  • Have deep intuition for modern generative AI architectures and how to squeeze performance out of them at inference time.
  • Are familiar with the internals of open-source inference frameworks (vLLM, SGLang, TensorRT-LLM, etc.) and can extend or replace them when needed.
  • Enjoy pathfinding new use cases — exploring heterogeneous deployment topologies and building early-stage POCs that prove out new ideas.
  • Are results-oriented with a strong bias toward action; you own problems end-to-end from prototype to optimization to handoff.
  • Are energized by working at the intersection of novel hardware and frontier models, and want your work to directly influence how next-generation AI silicon is used.
  • Value clear communication and thrive in a small, high-ownership team environment.
Responsibilities
  • Identify and prototype emerging LLM inference use cases suited to heterogeneous hardware deployments.
  • Build compelling proof-of-concept systems that demonstrate D-Matrix capabilities to customers, partners, and internal stakeholders.
  • Develop and tune custom kernels and operator-level optimizations to maximize throughput and minimize latency.
  • Drive quantization, sparsity, and batching strategies tailored to D-Matrix computational model.
  • Build and maintain inference runtimes, serving frameworks, and evaluation tooling.
  • Contribute to distributed inference systems: tensor/pipeline parallelism, disaggregated prefill/decode, KV-cache management.
  • Work closely with hardware architects to provide firmware and compiler teams with actionable inference workload insights.
  • Partner with product and business development to translate POCs into customer-facing demonstrations.
  • Contribute to technical publications, whitepapers, and open-source projects that advance D-Matrix visibility.
Required Qualifications
  • Bachelor’s degree in Computer Science, Electrical Engineering, or a related field, and 10+ years of relevant engineering experience; or equivalent demonstrated experience.
  • Master’s or PhD in Computer Science, Electrical Engineering, or a related field preferred, with 6+ years of relevant industry experience.
  • Strong proficiency in Python and C/C++.
  • Hands‑on experience optimizing LLM inference — attention kernels, KV cache, batching strategies, quantization (INT8/FP8/INT4).
  • Experience with at least one major inference framework (vLLM, SGLang, TensorRT-LLM, ONNX Runtime, or similar) at a contributor level.
  • Familiarity with GPU kernel programming (CUDA/Triton) and performance profiling tools.
Preferred Qualifications
  • Experience with heterogeneous compute deployments — scheduling inference workloads across dissimilar hardware (accelerators, CPUs, GPUs).
  • Familiarity with custom silicon or ASIC-based inference (beyond GPU-only environments).
  • Experience with distributed inference: tensor parallelism, pipeline parallelism, disaggregated serving.
  • Contributions to open-source inference or ML systems projects.
  • Experience with production inference serving at scale (latency SLOs, continuous batching, multi-model serving).
  • Familiarity with speculative decoding, mixture-of-experts routing, or long-context serving techniques.
  • Working familiarity with the material in the JAX Scaling Book or equivalent systems-level understanding of modern LLM training and inference.
Why D-Matrix Frontier Group
  • Work on genuinely novel hardware — D-Matrix in-memory compute architecture opens up inference optimization problems that don’t exist anywhere else.
  • End-to-end ownership from idea to deployed system, with a short feedback loop between your work and real hardware.
  • Small, senior team with high autonomy and direct influence on product direction.
  • Competitive compensation, equity, and benefits in Santa Clara, CA.
Equal Opportunity Employment Policy

d-Matrix is proud to be an equal opportunity workplace and affirmative action employer. We’re committed to fostering an inclusive environment where everyone feels welcomed and empowered to do their best work. We hire the best talent for our teams, regardless of race, religion, color, age, disability, sex, gender identity, sexual orientation, ancestry, genetic information, marital status, national origin, political affiliation, or veteran status. Our focus is on hiring teammates with humble expertise, kindness, dedication and a willingness to embrace challenges and learn together every day.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Staff LLM Inference Engineer
Senior Staff LLM Inference Engineer

Entrada Ventures • Santa Clara (CA)

On-site
USD 180,000 - 280,000
Competitive compensation
Equity opportunities
Benefits in Santa Clara, CA
Principal System Software Engineer, AI Inference Execution
Principal System Software Engineer, AI Inference Execution

d-Matrix • United States

Hybrid
USD 180,000 - 320,000
Principal System Software Engineer, AI Inference Execution
Principal System Software Engineer, AI Inference Execution

Entrada Ventures • Santa Clara (CA)

Hybrid
USD 180,000 - 260,000
Senior LLM Inference Architect — Heterogeneous Hardware
Senior LLM Inference Architect — Heterogeneous Hardware

d-Matrix inc. • Santa Clara (CA)

On-site
USD 130,000 - 170,000
Competitive compensation
Equity
Inclusive work environment
Senior LLM Inference Engineer - End-to-End, Equity
Senior LLM Inference Engineer - End-to-End, Equity

Entrada Ventures • Santa Clara (CA)

On-site
USD 180,000 - 280,000
Competitive compensation
Equity opportunities
Benefits in Santa Clara, CA
Sr. Staff, ML Researcher - LLM Algorithmic Optimization
Sr. Staff, ML Researcher - LLM Algorithmic Optimization

Entrada Ventures • Santa Clara (CA)

Hybrid
USD 180,000 - 240,000
Senior Staff ML Researcher - LLM Algorithmic Optimization
Senior Staff ML Researcher - LLM Algorithmic Optimization

d-Matrix inc. • Santa Clara (CA)

Hybrid
USD 130,000 - 160,000
Principal Architect, Performance Analysis and Modeling
Principal Architect, Performance Analysis and Modeling

D-Matrix Corp. • Santa Clara (CA), Northern (KY)

Hybrid
USD 150,000 - 230,000
Machine Learning Engineer, LLM Inference Optimization in Sonoma
Machine Learning Engineer, LLM Inference Optimization in Sonoma

NLP PEOPLE • Sonoma (CA)

On-site
USD 120,000 - 160,000
Principal Architect, Performance Analysis and Modeling
Principal Architect, Performance Analysis and Modeling

d-Matrix inc. • Santa Clara (CA)

Hybrid
USD 150,000 - 200,000