Get more replies from employers
Send a job-specific resume in minutes.
Adaption is seeking an experienced engineer to own the cost and performance of our inference stack in a production environment. You will influence batching, caching, quantization, and kernel-level optimization while coordinating with the serving fleet to meet throughput and latency targets.
You will optimize prefill/decode workloads, route performance between internal and external providers, and build profiling tools.
Adaption is seeking an experienced engineer to own the cost and performance of our inference stack in a production environment. You will influence batching, caching, quantization, and kernel-level optimization while coordinating with the serving fleet to meet throughput and latency targets.
You will optimize prefill/decode workloads, route performance between internal and external providers, and build profiling tools.