An application made for this job — a tailored resume and cover letter that speak straight to the posting.
Adaption Labs in San Francisco is looking for a senior ML systems engineer to own the cost and performance of the inference stack. You’ll optimize caching, batching, quantization, and kernel-level performance to improve throughput and reduce latency without sacrificing model quality.
You’ll collaborate with the serving fleet engineers and tune routing, profiling, and measurement systems across production workloads. A strong background in Python/C++/Rust and GPU optimization is required.
Adaption Labs in San Francisco is looking for a senior ML systems engineer to own the cost and performance of the inference stack. You’ll optimize caching, batching, quantization, and kernel-level performance to improve throughput and reduce latency without sacrificing model quality.
You’ll collaborate with the serving fleet engineers and tune routing, profiling, and measurement systems across production workloads. A strong background in Python/C++/Rust and GPU optimization is required.