Get more replies from employers
Send a job-specific resume in minutes.
Adaption is seeking an senior ML systems engineer to own the cost and performance of the inference stack. You will optimize caching, batching, and decoding while coordinating with the serving fleet to improve throughput and latency without compromising model quality.
You will work with engines like vLLM, SGLang, and TensorRT-LLM, implement profiling tools, and help route workloads to balance cost and performance in a global deployment.
Adaption is seeking an senior ML systems engineer to own the cost and performance of the inference stack. You will optimize caching, batching, and decoding while coordinating with the serving fleet to improve throughput and latency without compromising model quality.
You will work with engines like vLLM, SGLang, and TensorRT-LLM, implement profiling tools, and help route workloads to balance cost and performance in a global deployment.