ML Application Engineer – AI Inference & Model Optimization (Staff/Senior Staff level) - Riyadh, KSA

Qualcomm

As Sudiyah

On-site

OMR 30,760 - 46,140

Full time

14 days+
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Benefits offered by this job

Stock (RSU's) and performance-related bonus
Life and Medical Insurance
Relocation and immigration support

Job summary

Qualcomm is looking for a Machine Learning Applications Engineer – AI Inference & Model Optimization in Al Buraymi Governorate, Oman. This role involves optimizing deep learning workloads on Qualcomm's AI inference accelerators, focusing on model performance and embedded application support. The ideal candidate will have strong expertise in AI models and collaborate closely with clients to meet their technical needs.

The position offers an exciting opportunity to work at the intersection of AI silicon and system architecture while driving innovative solutions in a customer-facing role.

Qualifications

  • 10–15+ years of experience in deep learning model development or deployment.
  • Strong experience with model quantization and optimization techniques.
  • Proven ability to analyze model performance in production environments.

Responsibilities

  • Deploy and optimize deep learning models onto accelerator-based platforms.
  • Engage in hardware sizing and architecture discussions.
  • Assess AI model requirements and recommend alternative approaches.

Skills

AI model optimization
Deep learning model development
C/C++/Python programming
Collaboration with cross-functional teams
Machine learning frameworks

Education

Bachelor's degree in Computer Science or related field
Master's degree in Engineering or related field
PhD in Engineering or related field

Tools

TensorFlow
PyTorch
ONNX
Linux

Job description

Company

Qualcomm Middle East Information Technology Company LLC

Job Area

Engineering Group, Engineering Group > Software Engineering

General Summary
About Us

Qualcomm is enabling a world where everyone and everything can be intelligently connected. You interact with products and technologies made possible by Qualcomm every day, including intelligent edge devices, next-generation computing platforms, and advanced AI solutions. Qualcomm's leadership in AI, highperformance compute, and connectivity is driving innovation across cloud, edge, and data center environments - delivering scalable, powerefficient platforms that power the next generation of intelligent infrastructure.

About The Role

Qualcomm is seeking Machine Learning Applications Engineer – AI Inference & Model Optimization to support the enablement of rack-scale deep learning workloads on advanced Qualcomm AI inference accelerators. These accelerators utilize Qualcomm's expertise in hardware-accelerated AI to deliver high-performance, energy-efficient generative AI and computer vision inference solutions for modern data centers. This is a customer facing, highly technical role focused on porting, optimizing, and validating deep learning AI models on production systems, and enabling Qualcomm's partners to develop and deploy advanced machine learning applications - including computer vision, speech, generative AI and state of the art multimodal reasoning models - using popular frameworks such as PyTorch, TensorFlow, and ONNX on Qualcomm Cloud AI accelerators. Key responsibilities include evaluating models for throughput, latency, and accuracy; profiling and optimizing model performance; building robust application pipelines; integrating customer frameworks; and contributing to documentation, training, and demonstrations.

The role requires strong expertise in AI models, quantization, performance optimization, and deployment, plus the ability to shape architecture, workload sizing, and system design. It also requires experience with deep learning model development across hardware platforms, solid programming skills, collaboration with cross-functional teams, and proficiency in machine learning frameworks, Linux, and container orchestration tools.

The ideal candidate can effectively bridge AI model requirements hardware capabilities customer expectations, guiding customers from model selection hardware sizing deployment decisions production readiness.

What You'll Do
  • AI Model Porting & Optimization
  • Deploy, optimize and scale deep learning AI models onto acceleratorbased data center platforms, including:
  • Model conversion workflows
  • Quantization techniques (INT8 / mixed precision)
  • Runtime integration and optimization
  • Integrate ML models onto Qualcomm's Cloud AI ML stack from frameworks such as PyTorch, TensorFlow, and ONNX.
  • Drive improvements in model throughput, latency, and accuracy, with clear tradeoff analysis.
  • Build, test, and deploy scalable inference pipelines using serving frameworks such as vLLM, TGI, and Triton.
  • Optimize workloads for LLM and GenAI models across both multi-SoC and multi-card architectures.
  • Collaborate with engineering teams to analyze and refine training and inference for advanced deep learning applications.
  • Identify bottlenecks across compute, memory, and runtime, and guide optimization strategies.
  • Contribute to Qualcomm's Cloud AI GitHub repository and developer documentation, sharing technical best practices and solutions.
  • Develop and integrate end-to-end ML application pipelines with customer frameworks and libraries.
  • CustomerFacing Technical Engagement
  • Act as a trusted technical advisor for customers deploying AI workloads.
  • Engage in hardware sizing and architecture discussions, aligning model requirements with infrastructure capabilities.
  • Provide technical guidance on:
  • AI model selection
  • Deployment feasibility
  • System architecture and performance expectations
  • Lead discussions on model capabilities and limitations based on real customer use cases.
  • Model–Infrastructure Alignment
  • Assess and evaluate AI model requirements and recommend alternative model approaches when necessary.
  • Align model characteristics (latency, throughput, accuracy) with accelerator and system capabilities.
  • Connect model requirements with:
  • Memory constraints
  • Accelerator architecture
  • Scaling limitations
  • Support customers in defining model selection strategies based on deployment realities.
  • Performance & Scalability Engineering
  • Evaluate performance characteristics of AI models in production scenarios, including:
  • Throughput expectations
  • Latency targets
  • Concurrency behavior
  • Guide architecture decisions around:
  • Scaling strategies (horizontal vs vertical)
  • Hardware deployment sizing
  • Contribute to discussions on:
  • Workload scalability limits
  • Impact of model selection on system performance and efficiency
  • Provide insights into capacity planning and infrastructure optimization.
  • EndtoEnd AI Pipeline Design
  • Drive discussions around endtoend AI pipelines, including:
  • Multimodel workflows (e.g., detection + tracking + recognition)
  • Data preprocessing and postprocessing stages
  • Guide decisions on video and data processing stacks, including:
  • Video pipeline choices (e.g., FFMPEG vs GStreamer)
  • Integration into inference pipelines
  • Ensure pipelines are aligned with:
  • Performance requirements
  • Hardware capabilities
  • Realtime constraints
  • Model Tradeoff Analysis & Validation
  • Highlight and explain tradeoffs between:
  • Accuracy vs compatibility
  • Model quality vs deployment feasibility
  • Support decisionmaking on:
  • Model simplification vs performance gains
  • Precision vs efficiency tradeoffs
  • Lead or support model capability validation in deployment environments.
  • Collaborate with customers to define:
  • Inference assumptions
  • Model sizing strategies for largescale workloads
Required Qualifications
  • Bachelor's degree in Computer Science, Computer Engineering, Electrical Engineering, or related field (or equivalent experience).
  • 10–15+ years of experience in:
  • Deep learning model development or deployment experience on CPUs/GPUs/ASICs.
  • Inference systems and optimization
  • Data center or edge AI platforms
  • Strong experience with:
  • Model quantization and optimization techniques
  • AI model frameworks (e.g., PyTorch, TensorFlow)
  • Model deployment pipelines
  • Excellent C/C++/Python programming and software design skills, including debugging, and performance analysis.
  • Hands on expertise with Linux-based systems, low level software, drivers, and system bring up.
  • Proven ability to analyze and optimize model performance in production environments.
  • Solid understanding of:
  • AI inference hardware constraints
  • System level performance bottlenecks
  • Strong communication skills and experience in customer facing technical roles.
  • Willingness to travel for customer engagements and strategic reviews.
Preferred Qualifications
  • Skilled in deploying models on platforms that use hardware accelerators for inference.
  • Experienced with managing multi-model workflows and building real-time AI systems, including computer vision, video, and analytics projects.
  • Knowledgeable about distributed inference methods and handling large-scale model deployments.
  • Proficient in developing and maintaining video processing workflows and using relevant software frameworks.
  • Deep understanding of how system-level decisions affect performance in actual deployment environments.
  • Capable of simplifying complex technical ideas into straightforward, useful advice for clients.
  • Hands‑on experience running deep learning models on popular ML frameworks such as PyTorch, TensorFlow, ONNX
  • Experience developing software solutions that run in Linux environments with containers and orchestration
  • Experience with Source code and configuration management tools, Git knowledge is required.
  • Customer‑facing experience translating customer requirements into technical solutions (discovery, scoping, success criteria, and execution plans).
  • Proven ability to build and deliver technical demos, proofs‑of‑concept, and reference applications for ML/GenAI workloads.
  • Strong technical writing skills to produce customer‑ready documentation (getting started guides, deployment runbooks, troubleshooting guides) and deliver partner training sessions.
  • Experience driving issue triage and technical escalations with customers, coordinating across product, hardware, and software engineering teams to resolution.
  • Excellent stakeholder management and communication skills: present complex technical concepts clearly to both engineering and non‑engineering audiences.
Why Join Qualcomm

At Qualcomm, you'll work at the intersection of AI silicon, system architecture, and real world deployment. You will engage directly with strategic customers, influence nextgeneration AI data center platforms, and help define scalable, powerefficient infrastructure for the AI era. This role provides a unique opportunity to shape both technology direction and customer outcomes, while working with worldclass engineering and product teams.

What's On Offer
  • Salary including housing & transport allowance
  • Stock (RSU's) and performance related bonus
  • 16 weeks fully paid Maternity Leave
  • 6 weeks fully paid Paternity Leave
  • Employee stock purchase scheme
  • Child Education Allowance
  • Relocation and immigration support (if needed)
  • Life and Medical Insurance
  • Live+ Well Reimbursement for health and recreational membership fees
Minimum Qualifications
  • Bachelor's degree in Engineering, Information Systems, Computer Science, or related field and 6+ years of Software Engineering or related work experience.
  • Master's degree in Engineering, Information Systems, Computer Science, or related field and 5+ years of Software Engineering or related work experience.
  • PhD in Engineering, Information Systems, Computer Science, or related field and 4+ years of Software Engineering or related work experience.
  • 3+ years of work experience with Programming Language such as C, C++, Java, Python, etc.
  • References to a particular number of years experience are for indicative purposes only. Applications from candidates with equivalent experience will be considered, provided that the candidate can demonstrate an ability to fulfill the principal duties of the role and possesses the required competencies.

Qualcomm is an equal opportunity employer. If you are an individual with a disability and need an accommodation during the application/hiring process, rest assured that Qualcomm is committed to providing an accessible process.

Qualcomm expects its employees to abide by all applicable policies and procedures, including but not limited to security and other requirements regarding protection of Company confidential information and other confidential and/or proprietary information, to the extent those requirements are permissible under applicable law.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

AI Platform & Inference Suite Engineer (Staff/Senior Staff level) - Riyadh, KSA
AI Platform & Inference Suite Engineer (Staff/Senior Staff level) - Riyadh, KSA

Qualcomm • As Sudiyah

On-site
OMR 19,000 - 27,000
Housing and transport allowance
Stock options and performance bonus
Paid maternity and paternity leave
+2
Senior AI Inference Platform Engineer
Senior AI Inference Platform Engineer

Qualcomm • As Sudiyah

On-site
OMR 19,000 - 27,000
Housing and transport allowance
Stock options and performance bonus
Paid maternity and paternity leave
+2
Senior AI Inference & Model Optimization Engineer
Senior AI Inference & Model Optimization Engineer

Qualcomm • As Sudiyah

On-site
OMR 30,000 - 47,000
Stock (RSU's) and performance-related bonus
Life and Medical Insurance
Relocation and immigration support
Senior Staff ML Engineer: NPU-Optimized AI Frameworks
Senior Staff ML Engineer: NPU-Optimized AI Frameworks

Qualcomm • As Sudiyah

On-site
OMR 49,000 - 62,000
Salary including housing and transport allowance
Stock (RSU's) and performance-related bonus
16 weeks fully paid maternity leave
+6
Senior ML Engineer
Senior ML Engineer

NTG • Muscat

On-site
OMR 30,000 - 50,000
Senior Machine Learning Engineer / Technical Lead
Senior Machine Learning Engineer / Technical Lead

PhazeRo • Oman

On-site
OMR 40,000 - 55,000
Senior Machine Learning Engineer
Senior Machine Learning Engineer

Employment • Muscat

On-site
OMR 58,000 - 73,000
AI Engineer
AI Engineer

Esbaar • Muscat

On-site
OMR 12,000 - 28,000
Forward Deployed Engineer
Forward Deployed Engineer

Kore.ai • As Sudiyah

On-site
OMR 38,000 - 58,000
Assistant Manager/Manager - Next Gen-AI - Riyadh, Saudi Arabia
Assistant Manager/Manager - Next Gen-AI - Riyadh, Saudi Arabia

Protiviti India • As Sudiyah

On-site
OMR 38,000 - 50,000