Member of Technical Staff, Developer Relations

Inferact

United States

Remote

USD 200,000 - 400,000

Full time

14 days+
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Benefits offered by this job

Health benefits
Dental benefits
Vision benefits
401(k) company match

Job summary

Inferact seeks a Developer Relations Engineer to teach and evangelize vLLM as the default path for AI inference engineering. You will craft deep dives, demos, tutorials, and public docs to help practitioners build scalable systems.

You will cover topics like KV cache, continuous batching, quantization, and latency-throughput tradeoffs, while hosting workshops and writing accessible guides for the AI infrastructure community.

Qualifications

  • Bachelor’s degree or equivalent in CS/engineering/ML/systems.
  • Strong understanding of LLM inference systems and model serving.
  • Ability to credibly explain systems concepts (KV cache, PagedAttention, batching, prefill/decode, prefix caching).
  • Experience with vLLM or adjacent inference tech (SGLang, TensorRT-LLM, TGI, LoRAX, Ray Serve, FlashInfer, BentoML, Baseten-style).
  • Public portfolio of technical artifacts (blogs, tutorials, OSS docs, demos).
  • Ability to write and teach for practitioners without marketing tone.
  • Strong engineering judgment and ability to turn raw material into developer education.

Responsibilities

  • Write technical deep dives, build demos, tutorials, and contribute to docs and examples.
  • Host workshops and help developers understand topics like KV cache, continuous batching, and latency vs throughput.
  • Create public artifacts to help practitioners build better inference systems with vLLM.
  • Shape how the AI infrastructure community learns, adopts, and builds with vLLM.

Skills

LLM inference
Model serving
Distributed runtimes
Batching scheduling
Quantization
Tech education
Public portfolio
Engineering judgment
Clear communication

Education

Bachelor’s degree in CS/Engineering/ML

Tools

vLLM
TensorRT-LLM
Ray Serve
BentoML
PyTorch ecosystem

Job description

Overview

Inferact’s mission is to grow vLLM as the world’s AI inference engine and accelerate AI progress by making inference cheaper and faster. Founded by the creators and core maintainers of vLLM, we sit at the intersection of models and hardware, a position that took years to build.

About the Role

We’re looking for a Developer Relations Engineer to help make vLLM the default way developers understand, build, and scale AI inference. This is not a generic DevRel role. We’re looking for a inference systems educator-builder: someone who can understand vLLM as a deep LLM inference systems project, teach hard technical concepts clearly, and create public artifacts that help practitioners build better systems.

You’ll write technical deep dives, build demos, create tutorials, contribute to docs and examples, host workshops, and help developers understand topics like KV cache, continuous batching, prefix caching, prefill and decode, quantization, GPU serving, latency versus throughput, and model-server tradeoffs across vLLM and adjacent systems. Your work will shape how the broader AI infrastructure community learns, adopts, and builds with vLLM.

Skills and Qualifications

Minimum qualifications:

  • Bachelor’s degree or equivalent experience in computer science, engineering, machine learning, systems, or similar.

  • Strong technical understanding of LLM inference systems, model serving, GPU inference, distributed runtimes, scheduling, batching, quantization, or related infrastructure.

  • Ability to credibly explain systems concepts such as KV cache, PagedAttention, continuous batching, prefill / decode scheduling, prefix caching, speculative decoding, tensor parallelism, data parallelism, or latency versus throughput tradeoffs.

  • Experience with vLLM or adjacent inference technologies such as SGLang, TensorRT-LLM, TGI, LoRAX, Ray Serve, FlashInfer, BentoML, Baseten-style serving platforms, or similar systems.

  • A strong public portfolio of technical artifacts, such as blogs, tutorials, workshops, courses, OSS docs, benchmark posts, architecture explainers, conference talks, demos, or runnable repositories.

  • Ability to write and teach for practitioners without sounding like a content marketer.

  • Strong engineering judgment, product taste, and ability to turn raw technical material into useful developer education.

Preferred qualifications:

  • Prior work in ML systems, distributed systems, HPC, compilers, GPU kernels, serving infrastructure, MLOps, developer tooling, or open-source infrastructure.

  • Experience creating technical content that teaches reusable mental models, not just product features.

  • Experience contributing to developer-facing open source through docs, tutorials, examples, cookbooks, demos, or community support.

  • Existing credibility or community presence in AI infrastructure, OSS, CUDA / GPU, Ray, vLLM, PyTorch, Modal, BentoML, Baseten, Predibase, Together AI, Anyscale, LMSYS, or similar ecosystems.

  • Ability to host workshops, create hands-on labs, present technical talks, and help developers move from concept to working code.

Bonus points if you have:

  • Written widely-shared technical blogs, courses, or architecture deep dives on LLM inference, model serving, GPU serving, or ML systems.

  • Built demos, benchmarks, tutorials, or repositories around vLLM, SGLang, TensorRT-LLM, TGI, Ray Serve, FlashInfer, or related systems.

  • Contributed to open-source ML infrastructure, inference systems, developer tooling, or technical education projects.

  • Created practitioner-facing content with code, diagrams, benchmarks, demos, or end-to-end labs.

  • Built a durable personal portfolio that demonstrates technical depth, taste, and a strong point of view.

Logistics
  • Location: This role is based in San Francisco, California. Will consider remote in the US for exceptional candidates.

  • Compensation: Depending on background, skills, and experience, the expected annual salary range for this position is $200,000 - $400,000 USD + equity.

  • Visa sponsorship: We sponsor visas on a case-by-case basis.

  • Benefits: Inferact offers generous health, dental, and vision benefits as well as 401(k) company match.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Member of Technical Staff, Inference
Member of Technical Staff, Inference

Inferact • United States

Remote
USD 120,000 - 180,000
Visa sponsorship
Health coverage
Fully remote
Member of Technical Staff, Inference
Member of Technical Staff, Inference

Inferact • San Francisco (CA)

On-site
USD 200,000 - 400,000
Health, dental, and vision benefits
401(k) company match
Visa sponsorship on case-by-case basis
Head of Engineering
Head of Engineering

Inferact • San Francisco (CA)

On-site
USD 260,000 - 380,000
Health, dental, vision benefits
401(k) company match
Member of Technical Staff, Exceptional Generalist (Remote)
Member of Technical Staff, Exceptional Generalist (Remote)

Inferact • United States

On-site
USD 180,000 - 240,000
Competitive salary and equity
Visa sponsorship
Health coverage where applicable
Member of Technical Staff, Site Reliability Engineer
Member of Technical Staff, Site Reliability Engineer

Inferact • United States

Hybrid
USD 200,000 - 400,000
Member of Technical Staff, Performance and Scale
Member of Technical Staff, Performance and Scale

Inferact • San Francisco (CA)

On-site
USD 200,000 - 400,000
Generous health, dental, and vision benefits
401(k) company match
Equity options
Technical Product Marketing Manager
Technical Product Marketing Manager

Inferact • San Francisco (CA)

On-site
USD 140,000 - 210,000
Health, dental, and vision benefits
401(k) with company match
Member of Technical Staff, Cloud Orchestration (Remote)
Member of Technical Staff, Cloud Orchestration (Remote)

Inferact • United States

Remote
USD 140,000 - 190,000
Visa sponsorship
Fully remote
AI Inference Engineer
AI Inference Engineer

Premier Global Links • Palo Alto (CA)

On-site
USD 230,000 - 350,000
Developer Relations Engineer, AI Inference Systems
Developer Relations Engineer, AI Inference Systems

Inferact • United States

Remote
USD 200,000 - 400,000
Health benefits
Dental benefits
Vision benefits
+1