Sr Staff/Principal DevOps Engineer, Machine Learning Infrastructure and Performance

QUALCOMM, Inc.

Markham

On-site

CAD 120,000 - 165,000

Full time

3 days ago
Be an early applicant
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Job summary

Qualcomm Canada ULC is seeking an experienced ML Infrastructure Engineer to support on-site Linux systems and CI/CD pipelines. You will build and optimize GitHub Actions workflows, Docker environments, and a large model storage workflow.

You will also manage Python-based orchestration tools, performance data services (FastAPI/PostgreSQL), and plan scaling for distributed workloads, ensuring reliable builds, tests, and releases on hardware and edge devices.

Qualifications

  • Bachelor's, Master's, or PhD in a relevant field with multiple years of hardware, software, or systems engineering experience.
  • Experience with Linux-based infrastructure, CI/CD pipelines, and container tooling is required.
  • Strong programming and testing skills, especially Python, are essential.

Responsibilities

  • Operate and optimize a Linux self-hosted GitHub Actions runner fleet (capacity, scheduling, storage, monitoring, recovery).
  • Build reliable GitHub Actions pipelines and reproducible Docker environments for builds, tests, models, and releases.
  • Maintain Git LFS-backed model zoo and shared caches.
  • Own and improve Python-based orchestration and workflow tooling to ensure traceability and repeatability of hardware measurements.
  • Manage performance data ingestion/analysis services (FastAPI/PostgreSQL) and improve test reliability and release processes.
  • Plan scale steps, optimize caches and artifacts, and evaluate elastic or batch execution for large models.

Skills

Linux
CI/CD
Docker
Python

Education

Bachelor's degree in CS/Engineering/IS
Master's degree in CS/Engineering/IS
PhD in CS/Engineering/IS

Tools

GitHub Actions
Jenkins
Docker
Git LFS

Job description

Company:

Qualcomm Canada ULC

Job Area:

Engineering Group, Engineering Group > Machine Learning Engineering

General Summary:

Today, more intelligence is moving to end devices, and mobile is becoming the pervasive AI platform. Building on the smartphone foundation and the scale of mobile, Qualcomm envisions making AI ubiquitous—expanding beyond mobile and powering other end devices, machines, vehicles, and things.

Purpose:
  • Maintain local server racks and on-site devices
  • Make infrastructure and processes robust, reliable, and efficient
  • Identify and remedy items impacting productivity of the development team.
Responsibilities:
  • Operate and improve the Linux self-hosted GitHub runner fleet: capacity, scheduling, storage, monitoring, recovery, access, and incident response.
  • Build reliable GitHub Actions pipelines and reproducible Docker environments for builds, tests, model benchmarks, artifacts, and releases. Maintain the Git LFS-backed model zoo and its shared caches.
  • Own and innovate on the Python task/workflow orchestration infrastructure (Prefect-esque) to make hardware measurements traceable, repeatable, and actionable.
  • Steward performance-data ingestion and analysis, including the FastAPI/PostgreSQL-backed service and its clients; improve pytest integration, regression detection, reporting, triage, and release promotion.
  • Plan the next scale step: isolate workloads, improve cache and artifact lifecycle, and evaluate elastic/cloud or batch execution where it fits scarce devices and large models.
  • Turn project-specific tooling into supported, reusable platform components for other teams.
Required:
  • Depth in: Linux
  • Depth in: CI/CD
  • Depth in: Docker
  • Depth in: Python
The ideal candidate will be familiar with:
  • Linux infrastructure: Self-hosted GitHub Actions runners, systemd, remote filesystems (NFS), and resource monitoring.
  • CI/CD and containers: GitHub Actions, reusable workflows, Docker, and release automation. Familiarity with Jenkins is a plus.
  • Model and artifact management: Git LFS, shared caches, and large-model storage (ONNX models).
  • Python testing: pytest, pytest-xdist, integration tests, and performance reporting (with run-to-run variation).
  • AI performance tooling: ONNX, PyTorch, QAIRT SDK is a plus, Android device execution and profiling is a plus (adb).
  • Performance data: REST APIs, FastAPI, and PostgreSQL, or similar libraries/frameworks.
  • Future scaling: batch scheduling (ex. IBM Spectrum LSF) or cloud infrastructure for horizontally scaling automation.
Minimum Qualifications:
  • Minimum qualifications: Bachelor's degree in Computer Science, Engineering, Information Systems, or related field and 8+ years of Hardware Engineering, Software Engineering, Systems Engineering, or related work experience.
  • Minimum qualifications: Master's degree in Computer Science, Engineering, Information Systems, or related field and 7+ years of Hardware Engineering, Software Engineering, Systems Engineering, or related work experience.
  • Minimum qualifications: PhD in Computer Science, Engineering, Information Systems, or related field and 6+ years of Hardware Engineering, Software Engineering, Systems Engineering, or related work experience.
Equal Opportunity Statement

Applicants: Qualcomm is an equal opportunity employer. If you are an individual with a disability and need an accommodation during the application/hiring process, rest assured that Qualcomm is committed to providing an accessible process. You may e-mail disability-accomodations@qualcomm.com or call Qualcomm's toll-free number found here. Upon request, Qualcomm will provide reasonable accommodations to support individuals with disabilities to be able participate in the hiring process. Qualcomm is also committed to making our workplace accessible for individuals with disabilities. (Keep in mind that this email address is used to provide reasonable accommodations for individuals with disabilities. We will not respond here to requests for updates on applications or resume inquiries).

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Machine Learning Engineer, Frameworks & Performance (New Grad to Principal level)
Machine Learning Engineer, Frameworks & Performance (New Grad to Principal level)

QUALCOMM, Inc. • Markham

On-site
CAD 120,000 - 190,000
Machine Learning Engineer, Frameworks & Performance (New Grad to Principal level)
Machine Learning Engineer, Frameworks & Performance (New Grad to Principal level)

Qualcomm • Quebec

On-site
CAD 120,000 - 180,000
Sr Staff AI Software Engineer
Sr Staff AI Software Engineer

QUALCOMM, Inc. • Markham

On-site
CAD 110,000 - 170,000
Senior Engineer, AI Software Markham, Ontario, Canada Machine Learning Engineering Posted 20 hours ago
Senior Engineer, AI Software Markham, Ontario, Canada Machine Learning Engineering Posted 20 hours ago

Qualcomm • Markham

On-site
CAD 114,000 - 164,000
Sr. Software Engineer
Sr. Software Engineer

Qualcomm • Markham

On-site
CAD 80,000 - 110,000
Machine Learning Engineer, Frameworks & Performance (New Grad to Principal level)
Machine Learning Engineer, Frameworks & Performance (New Grad to Principal level)

Qualcomm • Montreal (administrative region)

On-site
CAD 120,000 - 170,000
Staff AI Software Engineer (C++)
Staff AI Software Engineer (C++)

Qualcomm • Markham

On-site
CAD 179,381 - 247,743
Competitive annual discretionary bonus
RSU grants
Comprehensive benefits package
Staff Software Engineer, Qualcomm AI Runtime
Staff Software Engineer, Qualcomm AI Runtime

Qualcomm • Markham

On-site
CAD 120,000 - 160,000
Sr. Engineer, AI Software/Machine Learning Engineering (C++)
Sr. Engineer, AI Software/Machine Learning Engineering (C++)

Qualcomm • Markham

On-site
CAD 157,923 - 226,945
Annual discretionary bonus
RSU grants
Competitive benefits package
Senior Embedded Software Engineer (C/C++), Machine Learning
Senior Embedded Software Engineer (C/C++), Machine Learning

Qualcomm • Markham

On-site
CAD 98,640 - 164,400