A complete application in a minute — tailored resume and cover letter, ready to send.
adaption is seeking an experienced engineer to own the cost and performance of our inference stack in the San Francisco Bay Area. You will influence throughput, latency, and reliability as workloads and hardware evolve.
You'll collaborate with the serving fleet engineers and own core levers like caching, batching, quantization, and decoding, ensuring efficient model serving without sacrificing quality.
adaption is seeking an experienced engineer to own the cost and performance of our inference stack in the San Francisco Bay Area. You will influence throughput, latency, and reliability as workloads and hardware evolve.
You'll collaborate with the serving fleet engineers and own core levers like caching, batching, quantization, and decoding, ensuring efficient model serving without sacrificing quality.