Senior Staff Infrastructure & Site Reliability Engineer – Datacentre AI Engineering - Riyadh, KSA Riyadh, Saudi Arabia Software Test Engineering Posted 16 hours ago

Qualcomm

Riyadh

On-site

SAR 450,000 - 650,000

Full time

14 days+
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Benefits offered by this job

Housing allowance
Transport allowance
Stock and performance bonus
Maternity/Paternity Leave

Job summary

Qualcomm Middle East Information Technology Company LLC is hiring Data Centre Engineers in Riyadh to design, operate, and continuously improve large-scale AI inference systems in a datacenter environment. You will ensure reliability, scalability, and production readiness for advanced ML workloads.

Ideal candidates have 12+ years in SRE, strong Python and Linux skills, IaC with Terraform/Ansible, and experience with CI/CD and monitoring tools in distributed systems.

Qualifications

  • Experience with AI/ML workloads and AI infrastructure.
  • Strong Python programming for production systems.
  • Experience with CI/CD pipelines and IaC tools.
  • Bachelor's/Master's in engineering or CS.

Responsibilities

  • Design and operate large-scale AI inference systems in datacenters.
  • Ensure reliability and scalability of AI clusters.
  • Develop tooling to support AI software stacks.
  • Collaborate with cross-functional teams across hardware and software.

Skills

Python
Linux
CI/CD
SRE Principles
Distributed Systems

Education

Bachelor’s or Master’s in Engineering/CS/AI

Tools

Terraform
Ansible
Prometheus
Grafana

Job description

Company

Qualcomm Middle East Information Technology Company LLC

Job Area

Engineering Group, Engineering Group > Software Test Engineering

General Summary

About Us

Qualcomm is growing its presence in Riyadh and is hiring Data Centre Engineers to support our expanding infrastructure across the region. As Saudi Arabia accelerates its digital transformation under Vision2030, Qualcomm is investing in world‑class computing and data centre capabilities to power AI, cloud, and advanced connectivity at scale. This is a unique opportunity to work in a fast‑growing technology hub, supporting critical environments and helping shape the future of data centre operations in the Kingdom and beyond.

About the Role

The role focuses on the design, operation, and continuous improvement of large‑scale AI inference systems in a datacenter environment. The engineer will support critical AI use cases by ensuring Qualcomm’s AI infrastructure is reliable, scalable, and production‑ready for advanced machine‑learning workloads.

The role requires strong systems and software engineering fundamentals, hands‑on execution, and the ability to work independently on complex problem areas while collaborating closely with cross‑functional teams across hardware, software, and machine learning.

Ideal candidate will have 12+ years of experience in SRE

Key Responsibilities
1. AI Infrastructure
  • Design, deploy, and operate large‑scale AI inference systems supporting critical AI workloads.
  • Ensure reliability, availability, and scalability of Qualcomm datacenter AI clusters.
  • Develop and maintain software tools and support infrastructure around AI software stacks.
2. AI & ML Engineering
  • Analyze software requirements and collaborate with architecture and hardware engineers to support AI workloads.
  • Build, deploy, and operate components supporting LLM inference, agentic AI workflows, and AI services.
  • Work with models, systems, and software teams to improve model performance on AI100 deployments.
  • Identify and implement optimizations for workloads running on multi‑SoC and multi‑card systems.
3. Site Reliability Engineering (SRE)
  • Apply SRE fundamentals including monitoring, alerting, incident response, and performance optimization.
  • Support production ML systems using MLOps tools and operational best practices.
  • Contribute to incident reviews, operational documentation, and continuous reliability improvements.
4. Observability & Tooling
  • Build and maintain observability tools, dashboards, and alerts to monitor system health and reliability.
  • Monitor infrastructure and services using tools such as Prometheus, Grafana, CloudWatch, and custom telemetry.
  • Create and maintain technical documentation, runbooks, and knowledge‑base articles.
5. Automation & CI/CD
  • Develop automation to reduce manual operational tasks and improve system reliability.
  • Support CI/CD pipelines for AI service and agent deployment.
  • Apply Infrastructure‑as‑Code practices using tools such as Terraform and Ansible.
Required Skillset
1. AI & Deep Learning
  • Experience working with AI/ML workloads such as LLMs, NLP, Vision, Audio, or Recommendation systems.
  • Understand ML inference concepts including batching, token streaming, and performance considerations.
  • Hands‑on experience with PyTorch and familiarity with modern ML frameworks.
  • Familiarity with distributed inference, checkpointing, and accelerator‑based compute environments.
2. AI Operations
  • Experience supporting AI or ML applications in production environments.
  • Familiarity with LLM inference pipelines and AI service operations.
3. Programming & Software Design
  • Strong programming skills in Python with experience building and supporting production systems.
  • Experience with scripting and automation using Python and Bash.
  • Familiarity with configuration management and orchestration tools.
4. Systems & Infrastructure
  • Strong Linux fundamentals include shell, containers, system services, and networking basics (DNS, TLS, HTTP/gRPC).
  • Experience working with cluster schedulers such as Slurm or equivalent systems.
  • Experience operating distributed systems with high availability and fault tolerance.
5. Observability & Monitoring
  • Hands‑on experience with monitoring and logging tools such as Prometheus, Grafana, ELK, or Loki.
  • Understanding of incident management, service health metrics, and system reliability monitoring.
6. DevOps & SRE Practices
  • Solid understanding of SDLC, release processes, and operational reliability practices.
  • Familiarity with CI/CD pipelines and Infrastructure‑as‑Code tools.
Preferred Skillset
  • Experience with GenAI, Agentic AI systems, or LLM orchestration frameworks.
  • Exposure to LangChain, AutoGen, or RAG‑based systems.
  • Experience with additional ML frameworks such as TensorFlow, JAX, or Ray.
  • Knowledge of GPU/accelerator‑based systems and high‑performance networking (RDMA, InfiniBand, RoCE).
  • Experience with advanced MLOps workflows or large‑scale AI platform operations.
Qualifications
  • Bachelor’s or Master’s degree in engineering, Computer Science, AI/ML, or a related field.
  • 12+ years of software, systems, or infrastructure engineering experience, preferably in production or datacenter environments.
What's on Offer
  • Salary including housing & transport allowance
  • Stock (RSU's) and performance related bonus
  • 16 weeks fully paid Maternity Leave
  • 6 weeks fully paid Paternity Leave
  • Employee stock purchase scheme
  • Child Education Allowance
  • Relocation and immigration support (if needed)
  • Life and Medical Insurance
  • Live+ Well Reimbursement for health and recreational membership fees
Minimum Qualifications
  • Bachelor's degree in Engineering, Information Systems, Computer Science, or related field and 6+ years of Software Test Engineering or related work experience.
  • Master's degree in Engineering, Information Systems, Computer Science, or related field and 5+ years of Software Test Engineering or related work experience.
  • PhD in Engineering, Information Systems, Computer Science, or related field and 4+ years of Software Test Engineering or related work experience.
  • 2+ year of work experience with Software Test or System Test, developing and automating test plans, and/or tools (e.g., Source Code Control Systems, Continuous Integration Tools, and Bug Tracking Tools).
  • *References to a particular number of years experience are for indicative purposes only. Applications from candidates with equivalent experience will be considered, provided that the candidate can demonstrate an ability to fulfill the principal duties of the role and possesses the required competencies.

Qualcomm is an equal opportunity employer. If you are an individual with a disability and need an accommodation during the application/hiring process, rest assured that Qualcomm is committed to providing an accessible process. You may e-mail disability-accomodations@qualcomm.com or call Qualcomm's toll-free number found here. Upon request, Qualcomm will provide reasonable accommodations to support individuals with disabilities to be able participate in the hiring process. Qualcomm is also committed to making our workplace accessible for individuals with disabilities. (Keep in mind that this email address is used to provide reasonable accommodations for individuals with disabilities. We will not respond here to requests for updates on applications or resume inquiries).

Qualcomm expects its employees to abide by all applicable policies and procedures, including but not limited to security and other requirements regarding protection of Company confidential information and other confidential and/or proprietary information, to the extent those requirements are permissible under applicable law.

If you would like more information about this role, please contact Qualcomm Careers.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Staff/Senior Staff Infrastructure & Site Reliability Engineer – Datacentre AI Engineering - Riyadh, KSA
Staff/Senior Staff Infrastructure & Site Reliability Engineer – Datacentre AI Engineering - Riyadh, KSA

Qualcomm • Riyadh

On-site
SAR 420,000 - 660,000
Relocation assistance
Immigration support
RSU's
+3
Staff/Senior Staff Infrastructure & Site Reliability Engineer – Datacentre AI Engineering - Riyadh, KSA
Staff/Senior Staff Infrastructure & Site Reliability Engineer – Datacentre AI Engineering - Riyadh, KSA

Qualcomm • Riyad Al Khabra

On-site
SAR 450,000 - 900,000
Housing & transport allowance
RSU stock & performance bonus
14–16 weeks maternity leave
+1
Riyadh, Saudi Arabia Software Engineering Posted 15 hours ago
Riyadh, Saudi Arabia Software Engineering Posted 15 hours ago

Qualcomm • Riyadh

On-site
SAR 250,000 - 420,000
Housing allowance
Performance bonus
Medical Insurance
Senior AI Infra & SRE Engineer - Datacenter
Senior AI Infra & SRE Engineer - Datacenter

Qualcomm • Riyadh

On-site
SAR 420,000 - 660,000
Relocation assistance
Immigration support
RSU's
+3
Senior AI Infra & SRE Engineer – Data Center
Senior AI Infra & SRE Engineer – Data Center

Qualcomm • Riyadh

On-site
SAR 450,000 - 650,000
Housing allowance
Transport allowance
Stock and performance bonus
+1
AI Infra & SRE Engineer - Data Center Scale
AI Infra & SRE Engineer - Data Center Scale

Qualcomm • Riyad Al Khabra

On-site
SAR 450,000 - 900,000
Housing & transport allowance
RSU stock & performance bonus
14–16 weeks maternity leave
+1
Test and Quality Lead
Test and Quality Lead

Think • Saudi Arabia

On-site
SAR 240,000 - 420,000
Competitive compensation
Early-stage ownership
Small senior team
+1
Middle+ General QA Engineer
Middle+ General QA Engineer

Jobgether • Saudi Arabia

On-site
SAR 120,000 - 180,000
Senior Hardware QA & Installation Support Engineer
Senior Hardware QA & Installation Support Engineer

Perfectvisionksa • Saudi Arabia

On-site
SAR 180,000 - 300,000
Senior Backend Engineer – AI Applications & OpenShift
Senior Backend Engineer – AI Applications & OpenShift

SoftwareONE Deutschland GmbH • Riyadh

Hybrid
SAR 320,000 - 520,000
Global company culture
Mentor support for start
President’s Club recognition
+2