Recevez plus de réponses des employeurs
Envoyez un CV adapté au poste en quelques minutes.
adaption is seeking an expert to own the cost and performance of its inference stack, driving throughput and latency improvements as workloads, traffic, and hardware evolve. You’ll partner with the serving fleet engineers to optimize caching, batching, quantization, and kernel-level tuning to achieve reliable, high-quality model deployments.
You will work with engines like vLLM, SGLang, and TensorRT-LLM while building profiling systems to illuminate time, memory, and compute usage across the
adaption is seeking an expert to own the cost and performance of its inference stack, driving throughput and latency improvements as workloads, traffic, and hardware evolve. You’ll partner with the serving fleet engineers to optimize caching, batching, quantization, and kernel-level tuning to achieve reliable, high-quality model deployments.
You will work with engines like vLLM, SGLang, and TensorRT-LLM while building profiling systems to illuminate time, memory, and compute usage across the