Recibe más respuestas de empleadores
Envía un currículum específico para el puesto de trabajo en cuestión de minutos.
Adaption is seeking a seasoned ML systems engineer to own the cost and performance of its inference stack. You will shape throughput and tail latency by tuning caching, batching, quantization, and decoding while preserving model quality.
You’ll collaborate with the serving fleet team, optimize routing to providers, and build profiling tools to reveal time, memory, and compute usage. Ideal candidates have 5+ years in ML systems and hands-on experience with vLLM, SGLang, or TensorRT-LLM, plus
Adaption is seeking a seasoned ML systems engineer to own the cost and performance of its inference stack. You will shape throughput and tail latency by tuning caching, batching, quantization, and decoding while preserving model quality.
You’ll collaborate with the serving fleet team, optimize routing to providers, and build profiling tools to reveal time, memory, and compute usage. Ideal candidates have 5+ years in ML systems and hands-on experience with vLLM, SGLang, or TensorRT-LLM, plus