Get more replies from employers
Send a job-specific resume in minutes.
Adaption is seeking an experienced ML systems engineer who will own the cost and performance of our inference stack. You will work with the serving fleet team to optimize caching, batching, quantization, and kernel-level tuning to maximize throughput and minimize tail latency without sacrificing model quality.
You will tune routing between internal infrastructure and external providers and build profiling tools to illuminate time and memory usage.
Adaption is seeking an experienced ML systems engineer who will own the cost and performance of our inference stack. You will work with the serving fleet team to optimize caching, batching, quantization, and kernel-level tuning to maximize throughput and minimize tail latency without sacrificing model quality.
You will tune routing between internal infrastructure and external providers and build profiling tools to illuminate time and memory usage.