Get more replies from employers
Send a job-specific resume in minutes.
Adaption Labs seeks an engineer to own the cost and performance of our inference stack. You will shape how we serve models as workloads shift, focusing on throughput, latency, and reliability.
You will collaborate with the serving fleet engineers, tuning caching, batching, quantization, decoding, and kernel-level optimizations to maximize efficiency without compromising model quality.
Adaption Labs seeks an engineer to own the cost and performance of our inference stack. You will shape how we serve models as workloads shift, focusing on throughput, latency, and reliability.
You will collaborate with the serving fleet engineers, tuning caching, batching, quantization, decoding, and kernel-level optimizations to maximize efficiency without compromising model quality.