Director, AI Inference & GPU-Accelerated Pipelines
WEKA
United States
On-site
USD 150,000 - 200,000
Full time
14 days+
Application generator
Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.
Get past ATS filters
Benefits offered by this job
Medical insurance
Dental insurance
Vision insurance
401(K) plan
Flexible Time off
Job summary
A growth-stage data infrastructure company is looking for a hands-on Director of Engineering - AI Inferences to lead a small team and architect high-performance AI inference systems. The ideal candidate will manage a team of developers, optimizing Large Language Model serving using frameworks like vLLM and LMCache. Expertise in backend engineering (Python, C++, or Rust) and experience with Kubernetes for scaling GPU workloads are essential. This role is perfect for someone eager to tackle complex data challenges in a fast-paced environment.
Qualifications
Proven experience with KV cache reuse, speculative decoding, and continuous batching.
Deep familiarity with vLLM, LMCache, and NIXL.
Expertise in GPU memory management.
Responsibilities
Architect and oversee the deployment of high-throughput, low-latency LLM inference pipelines.
Integrate and optimize serving engines to maximize hardware utilization.
Skills
Technical Leadership
Team Management
Inference Optimization
Backend Engineering
Kubernetes experience
Tools
Python
C++
Rust
CUDA
Job description
A growth-stage data infrastructure company is looking for a hands-on Director of Engineering - AI Inferences to lead a small team and architect high-performance AI inference systems. The ideal candidate will manage a team of developers, optimizing Large Language Model serving using frameworks like vLLM and LMCache. Expertise in backend engineering (Python, C++, or Rust) and experience with Kubernetes for scaling GPU workloads are essential. This role is perfect for someone eager to tackle complex data challenges in a fast-paced environment.