Get more replies from employers
Send a job-specific resume in minutes.
AMD is seeking a senior engineer to push AI inference performance on GPUs using vLLM as the primary serving framework. You will own hard performance problems across the stack, profiling, diagnosing, and optimizing models on customer-engaged deployments, delivering measurable uplifts and reusable methodologies.
You will work end-to-end across profiling, kernel and system-level optimizations, contributing to shared performance methodologies and accelerating daily work with AI agents.
AMD is seeking a senior engineer to push AI inference performance on GPUs using vLLM as the primary serving framework. You will own hard performance problems across the stack, profiling, diagnosing, and optimizing models on customer-engaged deployments, delivering measurable uplifts and reusable methodologies.
You will work end-to-end across profiling, kernel and system-level optimizations, contributing to shared performance methodologies and accelerating daily work with AI agents.