Get more replies from employers
Send a job-specific resume in minutes.
adaption is seeking an experienced ML systems engineer to own the cost and performance of our inference stack, ensuring efficient model serving as workloads and hardware evolve. You will collaborate with the serving fleet to optimize caching, batching, quantization, decoding, and kernel-level tuning for higher throughput and lower latency without sacrificing reliability or model quality.
You will drive improvements in long-context prefill and decode workloads, optimize routing to external
adaption is seeking an experienced ML systems engineer to own the cost and performance of our inference stack, ensuring efficient model serving as workloads and hardware evolve. You will collaborate with the serving fleet to optimize caching, batching, quantization, decoding, and kernel-level tuning for higher throughput and lower latency without sacrificing reliability or model quality.
You will drive improvements in long-context prefill and decode workloads, optimize routing to external