Director, AI Inference & GPU-Accelerated Pipelines

WEKA

United States

On-site

USD 150,000 - 200,000

Full time

14 days+
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Benefits offered by this job

Medical insurance
Dental insurance
Vision insurance
401(K) plan
Flexible Time off

Job summary

A growth-stage data infrastructure company is looking for a hands-on Director of Engineering - AI Inferences to lead a small team and architect high-performance AI inference systems. The ideal candidate will manage a team of developers, optimizing Large Language Model serving using frameworks like vLLM and LMCache. Expertise in backend engineering (Python, C++, or Rust) and experience with Kubernetes for scaling GPU workloads are essential. This role is perfect for someone eager to tackle complex data challenges in a fast-paced environment.

Qualifications

  • Proven experience with KV cache reuse, speculative decoding, and continuous batching.
  • Deep familiarity with vLLM, LMCache, and NIXL.
  • Expertise in GPU memory management.

Responsibilities

  • Architect and oversee the deployment of high-throughput, low-latency LLM inference pipelines.
  • Mentor and lead a team of developers.
  • Implement state-of-the-art KV cache management solutions.
  • Integrate and optimize serving engines to maximize hardware utilization.

Skills

Technical Leadership
Team Management
Inference Optimization
Backend Engineering
Kubernetes experience

Tools

Python
C++
Rust
CUDA

Job description

A growth-stage data infrastructure company is looking for a hands-on Director of Engineering - AI Inferences to lead a small team and architect high-performance AI inference systems. The ideal candidate will manage a team of developers, optimizing Large Language Model serving using frameworks like vLLM and LMCache. Expertise in backend engineering (Python, C++, or Rust) and experience with Kubernetes for scaling GPU workloads are essential. This role is perfect for someone eager to tackle complex data challenges in a fast-paced environment.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Tech Lead, Data & Inference Engineer: Scale AI Pipelines
Tech Lead, Data & Inference Engineer: Scale AI Pipelines

Catalyst Labs • New Jersey

Hybrid
USD 130,000 - 180,000
Tech Lead, Data & Inference Engineer: Scale AI Pipelines
Tech Lead, Data & Inference Engineer: Scale AI Pipelines

Catalyst Labs • Illinois

Hybrid
USD 120,000 - 160,000
Tech Lead, Data & Inference Engineer: Scale AI Pipelines
Tech Lead, Data & Inference Engineer: Scale AI Pipelines

Catalyst Labs • Palo Alto (CA)

On-site
USD 150,000 - 180,000
Lead Engineer, Inference Platform & Scale
Lead Engineer, Inference Platform & Scale

Cerebras • Sunnyvale (CA)

On-site
USD 150,000 - 200,000
Job stability with startup vitality
Open-source AI research
Simple, non-corporate work culture
Senior AI Inference Infrastructure Engineer
Senior AI Inference Infrastructure Engineer

Modular • United States

Hybrid
USD 167,000 - 273,000
Platform Lead, Data Pipelines & AI Inference
Platform Lead, Data Pipelines & AI Inference

Catalyst Labs • San Francisco (CA)

On-site
USD 180,000 - 280,000
Head of GPU-Accelerated AI Inference
Head of GPU-Accelerated AI Inference

NVIDIA • Washington

On-site
USD 224,000 - 432,000
Equity
Benefits package
Tech Lead, Data & Inference Engineer — AI Platform
Tech Lead, Data & Inference Engineer — AI Platform

Catalyst Labs • Manhattan Beach (CA)

Hybrid
USD 125,000 - 180,000
Engineering Manager, AI Inference — GPU-Accelerated DL
Engineering Manager, AI Inference — GPU-Accelerated DL

NVIDIA • California (MO)

On-site
USD 272,000 - 432,000
Staff Software Engineer, AI Inference
Staff Software Engineer, AI Inference

ChatGPT Jobs • New York (NY)

On-site
USD 180,000 - 240,000
Health Insurance
Equity