An application made for this job — a tailored resume and cover letter that speak straight to the posting.
Acceler8 Talent is seeking a Member of Technical Staff, Inference Performance to accelerate AI model inference across heterogeneous hardware. You will drive optimizations in latency, throughput, and memory usage, and push the serving stack with batching, caching, and quantization techniques.
The role involves developing CUDA/HIP/Triton kernels, deploying on accelerator clusters, and benchmarking models under real production workloads to improve efficiency and cost effectiveness.
Acceler8 Talent is seeking a Member of Technical Staff, Inference Performance to accelerate AI model inference across heterogeneous hardware. You will drive optimizations in latency, throughput, and memory usage, and push the serving stack with batching, caching, and quantization techniques.
The role involves developing CUDA/HIP/Triton kernels, deploying on accelerator clusters, and benchmarking models under real production workloads to improve efficiency and cost effectiveness.