Company: Qualcomm Technologies, Inc.
Job Area: Engineering Group > Machine Learning Engineering
General Summary: Qualcomm is utilizing its traditional strengths in digital wireless technologies to play a central role in the evolution of Cloud AI. The Qualcomm Cloud AI team is developing hardware and software solutions for Inference Acceleration.
Responsibilities
- Convert, optimize and deploy models for efficient inference using PyTorch and ONNX.
- Work at the forefront of GenAI by understanding advanced algorithms (e.g., attention mechanisms, MoEs) and numerics to identify new optimization opportunities.
- Analyze performance and optimize LLM, VLM, and diffusion models for inference, scaling performance for throughput and latency constraints.
- Map next-generation AI workloads on top of current and future hardware designs.
- Collaborate with customers and internal compiler, firmware, and platform teams to drive solutions.
- Analyze complex performance or stability issues to determine root causes.
- Create engineering solutions that deliver continuous insights into performance of AI workloads, guiding improvements over time.
- Design and implement high-level kernels (e.g., in Triton) focused on generating efficient, low-level code.
Qualifications
- Hands‑on experience building and optimizing language models, notably in PyTorch and ONNX, preferably in production‑grade environments.
- Deep understanding of transformer architectures, attention mechanisms, and performance trade‑offs.
- Experience with workload‑mapping strategies exhibiting sharding or various parallelisms.
- Strong Python programming skills.
- Proactive learning about the latest inference optimization techniques.
- Understanding of computer architecture, ML accelerators, in‑memory processing, and distributed systems.
- Strong communication, problem‑solving skills, and ability to work effectively in a fast‑paced, collaborative environment.
- MS in Computer Science, Machine Learning, Computer Engineering, or Electrical Engineering.
Bonus Skills
- Background in neural network operators and mathematical operations, including linear algebra and math libraries.
- Understanding of machine learning compilers.
- Experience in converging accuracy and its evaluation methods.
- Knowledge of torch.compile or torchDynamo.
- PhD in Computer Science, Computer Engineering, or Machine Learning.
Minimum Qualifications
- Bachelor's degree in Computer Science, Engineering, Information Systems, or related field and 6+ years of hardware, software, or systems engineering experience.
- Master's degree in Computer Science, Engineering, Information Systems, or related field and 5+ years of hardware, software, or systems engineering experience.
- PhD in Computer Science, Engineering, Information Systems, or related field and 4+ years of hardware, software, or systems engineering experience.
Compensation & Benefits
- Pay range: $178,400.00 – $267,600.00.
- Competitive annual discretionary bonus program and annual RSU grants.
- Highly competitive benefits package.
Qualcomm is an equal opportunity employer. If you are an individual with a disability and need accommodation during the application/hiring process, Qualcomm is committed to providing an accessible process and will provide reasonable accommodations to support participation.