ml engineer for LLM inference

Cast AI

United States

On-site

USD 180,000 - 270,000

Full time

4 days ago
Be an early applicant
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Benefits offered by this job

Equity 10%
Annual hackathon
Learning budget

Job summary

Cast AI ищет инженера по ML-инфраструктуре для масштабирования и оптимизации производственных ML-систем. Вы будете работать над инференсом в распределённых конфигурациях, экспериментировать с квантованием и ускорением на разных GPU, улучшать время отклика и пропускную способность в рамках гибридной облачной платформы.

Требуется опыт работы с инференс-инфраструктурой и сильные навыки Python, а также портфолио проектов с оптимизацией и профильными знаниями в vLLM/SGLang/TensorRT-LLM.

Qualifications

  • 5+ лет опыта в построении production ML-систем.
  • Портфолио с инференсом/инфраструктурой обучения, а не только ноутбуки.
  • Сильные навыки Python для production services.
  • Опыт с vLLM, SGLang или TensorRT-LLM и понимание производительности инференса на разных GPU.
  • Понимание компромиссов в квантовании и качественные регрессии.
  • Опыт в распределённых системах, шардинге и отказоустойчивости в многопроцессорных конфигурациях.
  • Ориентация на измерения: инструментация перед оптимизацией.

Responsibilities

  • Увеличение пропускной способности через батчинг, speculative decoding, chunked prefill и настройку ядра.
  • Снижение задержки за счёт профилирования узких мест и оптимизации вычислений, памяти и сетей.
  • Улучшение использования KV кэша через пагинацию внимания и кэширование.
  • Квантизация весов/активаций и KV (INT8/INT4/FP8) с контролем качества.
  • Снижение холодных стартов за счёт быстрой инициализации и более эффективной загрузки весов.
  • Масштабирование инференса по узлам, выбор топологий и checkpointing.
  • Определение технического направления: какие метрики, какие тесты и воспроизводимые эксперименты проводить.

Skills

Python
Production ML Systems
Distributed Systems
Inference Infrastructure
Quantization

Tools

vLLM
SGLang
TensorRT-LLM

Job description

Описание:

Cast AI is an automation platform that operates cloud-native and AI infrastructure at scale. It uses autonomous decision-making in Kubernetes and cloud environments to continuously optimize production performance, reliability, and efficiency. Its Kimchi system automatically matches workloads to cost-efficient, high-performing LLM and serving configurations on customer infrastructure.

Задачи:
  • Increase throughput through continuous batching, speculative decoding, chunked prefill, and kernel-level tuning across vLLM, SGLang, and TensorRT-LLM
  • Reduce TTFT and TPOT latency by profiling bottlenecks and addressing compute, memory bandwidth, scheduling, or networking issues
  • Improve KV cache utilization through paged attention, prefix caching, eviction policies, cache reuse, and quantized KV
  • Quantize weights, activations, and KV using INT8, INT4, and FP8 while measuring quality on real workloads
  • Reduce cold starts and memory footprint through faster initialization, smarter weight loading, and tighter memory accounting
  • Scale inference across nodes using distributed topologies, network-aware placement, and checkpointing strategies
  • Set the technical direction by deciding what to benchmark, adopt, and build, and align the team through writeups and reproducible experiments
Требования:
  • 5+ Years building production ML systems
  • A portfolio demonstrating depth in inference or training infrastructure, not just model training notebooks
  • Strong Python skills for production services
  • Hands-on experience with at least one of vLLM, SGLang, or TensorRT-LLM and an understanding of inference engine performance on different GPUs
  • Fluency with quantization tradeoffs, including measuring quality regressions
  • Experience with distributed systems, collective communication, sharding strategies, and failure modes in multi-GPU and multi-node setups
  • A measurement-focused approach: instrument before optimizing and distinguish real improvements from benchmark artifacts
  • Self-direction and comfort with a broad mandate
Условия:
  • Competitive salary depending on experience level
  • Equity options 10%
  • Of work time for personal projects or self-improvement
  • Learning budget for professional and personal development, including international conferences and courses
  • Annual hackathon
  • Team-building budget and company events
  • Equipment budget
  • Extra days off
  • No visa sponsorship or work permit is provided
  • A background check may be conducted at the final recruitment stage through Checkr
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

ml engineer for llm performance optimization
ml engineer for llm performance optimization

NDA • United States

On-site
USD 150,000 - 230,000
Leave + wellness days
Training coverage
Health stipend
+1
ml engineer for AI companions
ml engineer for AI companions

Joi AI • United States

Remote
USD 150,000 - 210,000
Ежегодный отпуск 28 дней
Ежегодные wellness-дни
Referral бонус до $5000
+3
ML engineer LLM inference & architecture
ML engineer LLM inference & architecture

NDA • United States

On-site
USD 180,000 - 240,000
Fully remote
Home office stipend
Education support
+4
nlp engineer for llm inference
nlp engineer for llm inference

NDA • United States

On-site
USD 120,000 - 180,000
Высокое вознаграждение
Возможность роста до Head of ML
ml engineer для инференса LLM
ml engineer для инференса LLM

NDA • United States

On-site
USD 140,000 - 240,000
Principal Software Engineer, Inference
Principal Software Engineer, Inference

Hewlett Packard Enterprise • Fort Collins (CO)

On-site
USD 180,000 - 250,000
Health and wellbeing benefits
Professional development programs
Flexible work arrangements
ai engineer for LLM applications
ai engineer for LLM applications

EPAM • United States

Remote
USD 140,000 - 200,000
ml engineer for scalable AI solutions
ml engineer for scalable AI solutions

Andersen • United States

On-site
USD 120,000 - 180,000
Annual bonus
Private health insurance
Relocation assistance
+8
AI Engineer / LLM Systems Engineer
AI Engineer / LLM Systems Engineer

Sphere Software • United States

Remote
USD 120,000 - 170,000
Software Engineer, ML Infrastructure
Software Engineer, ML Infrastructure

Realm Labs • Sunnyvale (CA)

On-site
USD 210,000 - 350,000
Market aligned compensation
Founding engineer equity
Medical, Dental, Vision, and Life insurance
+2