Get more replies from employers
Send a job-specific resume in minutes.
AMD is looking for a performance-obsessed engineer to drive AI inference performance to the absolute limit on AMD GPUs, with SGLang as the primary serving framework. You will lead a small, highly technical team and work end-to-end across the stack: profiling, diagnosing, and optimizing leading models running on SGLang across customer-relevant serving configurations (e.g.
agentic coding, long-context, high-throughput serving).
AMD is looking for a performance-obsessed engineer to drive AI inference performance to the absolute limit on AMD GPUs, with SGLang as the primary serving framework. You will lead a small, highly technical team and work end-to-end across the stack: profiling, diagnosing, and optimizing leading models running on SGLang across customer-relevant serving configurations (e.g. agentic coding, long-context, high-throughput serving).
Master's, or PhD in Computer Science, Computer Engineering, Electrical Engineering, or equivalent. 7+ years of software development experience in GPU computing, AI systems, or high-performance computing, Deep hands-on experience with SGLang internals, Strong background in end-to-end workload profiling and bottleneck diagnosis, Understanding of GPU kernel performance characteristics (occupancy, register/LDS pressure, memory coalescing, cache utilization), Understanding of model architectures (transformers, MoE, diffusion) and inference paradigms (speculative decoding, continuous batching), Experience with custom kernel development or integration (HIP, CUDA, Triton, CK, or similar), Understanding of multi-GPU and multi-node distributed systems (RCCL/NCCL, RDMA), Strong proficiency in Python and C++, Customer-facing technical leadership experience, Strong Linux systems knowledge, Excellent written and verbal English communication skills, Master's or PhD in Computer Science, Computer Engineering, Electrical Engineering, or equivalent