Get more replies from employers
Send a job-specific resume in minutes.
Adaption is seeking a senior ML systems engineer to own the cost and performance of our inference stack in a rapidly evolving environment. You will shape throughput, latency, and reliability by tuning caching, batching, and kernels while collaborating with the serving fleet engineers.
You will work with engines like vLLM, SGLang, and TensorRT-LLM, and build tools to measure where compute is spent. A strong background in Python and systems languages is essential, with a track record of real-world
Adaption is seeking a senior ML systems engineer to own the cost and performance of our inference stack in a rapidly evolving environment. You will shape throughput, latency, and reliability by tuning caching, batching, and kernels while collaborating with the serving fleet engineers.
You will work with engines like vLLM, SGLang, and TensorRT-LLM, and build tools to measure where compute is spent. A strong background in Python and systems languages is essential, with a track record of real-world