Company
Qualcomm Technologies, Inc.
Job Area
Engineering Group, Engineering Group > Machine Learning Engineering
Job Description
This is a full‑time onsite role requiring five days a week in the office at Qualcomm’s San Diego location. As a Staff or Senior Staff Software Engineer on the Qualcomm AI Stack SDK Software team, you will design, develop, and deliver advanced AI/ML software solutions for generative AI inference on Snapdragon platforms. The role focuses on model optimization, quantization, graph transformations, and runtime execution for modern AI architectures, including LLMs, LVMs, and LMMs. You will work at the intersection of machine learning algorithms, inference optimization, graph lowering, and systems software, contributing directly to the Qualcomm AI Stack SDK (QAIRT) and associated tools, such as delegates support for ONNX Runtime, ExecuTorch, and TFLite/LiteRT frameworks. Collaboration with engineers across multiple teams—ML Research, AI accelerator HW/SW, Product Management, Program Management, and QA—is central to driving features from concept to production. This position requires strong technical ownership, independent work, and the capability to deliver features end‑to‑end while mentoring junior engineers.
Responsibilities
- Convert, optimize, and deploy AI models from PyTorch and ONNX frameworks for efficient inference on Snapdragon platforms.
- Design and implement graph transformations, graph lowering, and optimization techniques within AI runtime environments such as ONNX Runtime, ExecuTorch, and Qualcomm AI Stack SDK.
- Apply knowledge of quantization and performance optimization to improve latency, throughput, memory usage, and power efficiency.
- Work at the forefront of generative AI, understanding advanced algorithms such as attention mechanisms, mixture‑of‑experts (MoE), low‑rank adapter (LoRA), and emerging inference optimization techniques (e.g., speculative decoding).
- Collaborate with ML Research teams to prototype and productize new features and techniques into SDK solutions.
- Debug complex issues across models, runtime, OS, compiler, and hardware layers, working closely with QA and customer teams.
- Design, implement, and deliver new features and enhancements to the Qualcomm AI Stack SDK.
- Participate in design reviews and code reviews, ensuring software quality and maintainability.
- Mentor junior engineers, helping them prioritize work and drive execution across multiple initiatives.
Minimum Qualifications
- Bachelor’s degree in Computer Science, Engineering, Information Systems, or related field and 4+ years of hardware engineering, software engineering, systems engineering, or related experience.
- Master’s degree in Computer Science, Engineering, Information Systems, or related field and 3+ years of hardware engineering, software engineering, systems engineering, or related experience.
- PhD in Computer Science, Engineering, Information Systems, or related field and 2+ years of hardware engineering, software engineering, systems engineering, or related experience.
- Bachelor’s degree in computer science, computer engineering, or related field and 6+ years (Staff) or 8+ years (Senior Staff) of experience in software design, development, and delivery.
- Master’s degree or PhD in computer science, computer engineering, or related field and 5+ years (Staff) or 7+ years (Senior Staff) of experience in software design, development, and delivery.
- At least 3+ years of hands‑on experience in AI/ML software development, focusing on inference or model optimization.
- Strong understanding of AI/ML fundamentals, including deep learning and inference pipelines.
- Deep understanding of transformer architectures, attention mechanisms, and performance trade‑offs.
- Proficiency in Python and C/C++ for production‑quality software development.
- Experience working with PyTorch and ONNX models and tooling.
- Debugging skill of complex issues, performing root cause analysis, and ensuring high system reliability.
- Ability to work independently, collaborate across teams, and drive complex features end‑to‑end.
Preferred Qualifications
- Working knowledge of graph theory, graph optimizations, and compiler‑style transformations.
- Experience with LLM, LVM, and LMM inference pipelines, including prefill and generation workflows.
- Familiarity with Hugging Face ecosystem, including model repositories and interfaces such as PEFT.
- Experience with LoRA, MoE‑based models, and awareness of modern GenAI inference techniques.
- Experience with Android and/or RTOS environments (e.g., QNX).
- Experience with CMake‑based build environments, agile software development practices, and git‑based SCM.
- At least 2 years of experience in embedded software or system‑level software development and optimization.
- At least 2 years of experience interacting with senior leadership (Director level and above).
- Ability to collaborate across a globally diverse team and manage multiple priorities.
- Previous experience mentoring junior engineers.
- User‑level or development experience with Qualcomm AI Stack/SDKs (e.g., QAIRT, QNN, Genie).
- Exposure to Snapdragon SoCs and AI accelerators such as NPU.
- Prior hands‑on experience with GenAI features such as transformer architectures, LoRA, MoE, speculative decoding, and vision encoder/decoder models.
Pay Range
$158,400.00 – $237,600.00
EEO Statement
Qualcomm is an equal opportunity employer. All qualified applicants will receive consideration for employment without regard to race, color, religion, sex, sexual orientation, gender identity, national origin, disability, veteran status, or any other protected classification.
Disability Accommodations
Qualcomm is committed to providing an accessible process for individuals with disabilities. Contact email-accommodations@qualcomm.com or call Qualcomm’s toll‑free number for assistance. Qualcomm will not respond to application updates or resume inquiries via this email address.