Sr Staff/Principal DevOps Engineer, Machine Learning Infrastructure and Performance

Linuxassociation

Markham

Hybrid

CAD 140,000 - 200,000

Full time

2 days ago
Be an early applicant
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Job summary

Qualcomm is seeking a Sr Staff/Principal DevOps Engineer to join the Machine Learning Infrastructure and Performance team in Markham, Ontario. You will maintain on-site servers and devices, improve CI/CD pipelines and model tooling, and plan scalable infrastructure for large ML workloads.

The role requires deep Linux expertise, CI/CD with GitHub Actions, containerization with Docker, and strong Python-based automation.

Qualifications

  • Bachelor's degree in CS/Engineering or related field with 8+ years of relevant work experience
  • Master's degree in CS/Engineering or related field with 7+ years of relevant work experience
  • PhD in CS/Engineering or related field with 6+ years of relevant work experience

Responsibilities

  • Operate and improve the Linux self-hosted GitHub runner fleet: capacity, scheduling, storage, monitoring, recovery, access, incident response.
  • Build reliable GitHub Actions pipelines and reproducible Docker environments for builds, tests, model benchmarks, artifacts, and releases.
  • Own and innovate on the Python task/workflow orchestration infrastructure (Prefect-esque).
  • Steward performance-data ingestion and analysis (FastAPI/PostgreSQL).
  • Plan the next scale step: isolate workloads, improve cache and artifact lifecycle, evaluate elastic/cloud or batch execution.
  • Turn tooling into reusable platform components for other teams.

Skills

Linux
CI/CD
Docker
Python
GitHub Actions
FastAPI
PostgreSQL
ONNX

Education

Bachelor's degree
Master's degree
PhD

Tools

Git LFS
Prefect-esque orchestration
ONNX models
PyTorch

Job description

# Sr Staff/Principal DevOps Engineer, Machine Learning Infrastructure and PerformanceQualcomm · Markham, Ontario, CABack to resultsSr Staff/Principal DevOps Engineer, Machine Learning Infrastructure and PerformanceQualcommMarkham, Ontario, CA Posted October 8, 2026 Via LinkedIn### Full Job DescriptionCompanyQualcomm Canada ULCJob AreaEngineering Group, Engineering Group > Machine Learning EngineeringGeneral SummaryToday, more intelligence is moving to end devices, and mobile is becoming the pervasive AI platform. Building on the smartphone foundation and the scale of mobile, Qualcomm envisions making AI ubiquitous—expanding beyond mobile and powering other end devices, machines, vehicles, and things.We are inventing, developing, and commercializing power-efficient on-device AI, edge cloud AI, and 5G to make this a reality. New PositionPurposeAs a member of Qualcomm’s ML Systems Team, you will:• Maintain local server racks and on-site devices• Make infrastructure and processes robust, reliable, and efficient• Identify and remedy items impacting productivity of the development team.Responsibilities• Operate and improve the Linux self-hosted GitHub runner fleet: capacity, scheduling, storage, monitoring, recovery, access, and incident response.• Build reliable GitHub Actions pipelines and reproducible Docker environments for builds, tests, model benchmarks, artifacts, and releases. Maintain the Git LFS-backed model zoo and its shared caches.• Own and innovate on the Python task/workflow orchestration infrastructure (Prefect-esque) to make hardware measurements traceable, repeatable, and actionable.• Steward performance-data ingestion and analysis, including the FastAPI/PostgreSQL-backed service and its clients; improve pytest integration, regression detection, reporting, triage, and release promotion.• Plan the next scale step: isolate workloads, improve cache and artifact lifecycle, and evaluate elastic/cloud or batch execution where it fits scarce devices and large models.• Turn project-specific tooling into supported, reusable platform components for other teams.RequiredDepth in:• Linux• CI/CD• Docker• PythonThe Ideal Candidate Will Be Familiar With• Linux infrastructure: Self-hosted GitHub Actions runners, systemd, remote filesystems (NFS), and resource monitoring.• CI/CD and containers: GitHub Actions, reusable workflows, Docker, and release automation. Familiarity with Jenkins is a plus.• Model and artifact management: Git LFS, shared caches, and large-model storage (ONNX models).• Python testing: pytest, pytest-xdist, integration tests, and performance reporting (with run-to-run variation).• AI performance tooling: ONNX, PyTorch, QAIRT SDK is a plus, Android device execution and profiling is a plus (adb).• Performance data: REST APIs, FastAPI, and PostgreSQL, or similar libraries/frameworks.• Future scaling: batch scheduling (ex. IBM Spectrum LSF) or cloud infrastructure for horizontally scaling automation.Minimum Qualifications• Bachelor's degree in Computer Science, Engineering, Information Systems, or related field and 8+ years of Hardware Engineering, Software Engineering, Systems Engineering, or related work experience.ORMaster's degree in Computer Science, Engineering, Information Systems, or related field and 7+ years of Hardware Engineering, Software Engineering, Systems Engineering, or related work experience.ORPhD in Computer Science, Engineering, Information Systems, or related field and 6+ years of Hardware Engineering, Software Engineering, Systems Engineering, or related work experience.Applicants: Qualcomm is an equal opportunity employer. If you are an individual with a disability and need an accommodation during the application/hiring process, rest assured that Qualcomm is committed to providing an accessible process. You may e-mail disability-accomodations@qualcomm.com or call Qualcomm's toll-free number found here. Upon request, Qualcomm will provide reasonable accommodations to support individuals with disabilities to be able participate in the hiring process. Qualcomm is also committed to making our workplace accessible for individuals with disabilities. (Keep in mind that this email address is used to provide reasonable accommodations for individuals with disabilities. We will not respond here to requests for updates on applications or resume inquiries).Qualcomm expects its employees to abide by all applicable policies and procedures, including but not limited to security and other requirements regarding protection of Company confidential information and other confidential and/or proprietary information, to the extent those requirements are permissible under applicable law.To all Staffing and Recruiting Agencies: Our Careers Site is only for individuals seeking a job at Qualcomm. Staffing and recruiting agencies and individuals being represented by an agency are not authorized to use this site or to submit profiles, applications or resumes, and any such submissions will be considered unsolicited. Qualcomm does not accept unsolicited resumes or applications from agencies. Please do not forward resumes to our jobs alias, Qualcomm employees or any other company location. Qualcomm is not responsible for any fees related to unsolicited resumes/applications.If you would like more information about this role, please contact Qualcomm Careers.
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Sr Staff/Principal DevOps Engineer, Machine Learning Infrastructure and Performance
Sr Staff/Principal DevOps Engineer, Machine Learning Infrastructure and Performance

QUALCOMM, Inc. • Markham

On-site
CAD 120,000 - 165,000
Machine Learning Engineer, Frameworks & Performance (New Grad to Principal level)
Machine Learning Engineer, Frameworks & Performance (New Grad to Principal level)

QUALCOMM, Inc. • Markham

On-site
CAD 120,000 - 190,000
Machine Learning Engineer, Frameworks & Performance (New Grad to Principal level)
Machine Learning Engineer, Frameworks & Performance (New Grad to Principal level)

Qualcomm • Quebec

On-site
CAD 120,000 - 180,000
Sr Staff AI Software Engineer
Sr Staff AI Software Engineer

QUALCOMM, Inc. • Markham

On-site
CAD 110,000 - 170,000
Senior Engineer, AI Software Markham, Ontario, Canada Machine Learning Engineering Posted 20 hours ago
Senior Engineer, AI Software Markham, Ontario, Canada Machine Learning Engineering Posted 20 hours ago

Qualcomm • Markham

On-site
CAD 114,000 - 164,000
Staff Software Engineer, Qualcomm AI Runtime
Staff Software Engineer, Qualcomm AI Runtime

Qualcomm • Markham

On-site
CAD 120,000 - 160,000
Engineer, AI Software/Machine Learning (C++)
Engineer, AI Software/Machine Learning (C++)

Qualcomm • Markham

On-site
CAD 105,000 - 155,000
Sr. Software Engineer
Sr. Software Engineer

Qualcomm • Markham

On-site
CAD 80,000 - 110,000
Engineer, AI Software/Machine Learning (C++)
Engineer, AI Software/Machine Learning (C++)

QUALCOMM, Inc. • Markham

On-site
CAD 100,000 - 149,000
Machine Learning Engineer, AI Processors (New Grad to Engineer Level)
Machine Learning Engineer, AI Processors (New Grad to Engineer Level)

QUALCOMM, Inc. • Markham

On-site
CAD 105,000 - 155,000
RSU grants
Bonus program
Benefits package