Staff DevOps Engineer — AI/ML Infra at Scale

Nexxa.ai

Canada

On-site

CAD 120,000 - 180,000

Full time

5 days ago
Be an early applicant
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Benefits offered by this job

Equity package
Professional growth opportunities

Job summary

Nexxa.ai in Canada is seeking a Senior/Staff DevOps Engineer to own and operate the infrastructure for AI and industrial systems at scale. You will oversee compute, networking, storage, and deployment, ensuring production ML workloads run fast, with strong observability and reliability.

You will collaborate with AI, data, and product teams, design CI/CD pipelines, IaC with Terraform/Pulumi, and manage Kubernetes platforms with GPU scheduling.

Qualifications

  • 6+ years in DevOps/SRE/Platform/infra roles.
  • Production-scale cloud platforms (AWS, GCP, or Azure)
  • Kubernetes in production with GPU scheduling
  • Infrastructure-as-code tooling (Terraform, Pulumi, or equivalent)
  • CI/CD systems (GitHub Actions, GitLab CI, CircleCI, Jenkins, ArgoCD)
  • Observability stacks (Prometheus, Grafana, Datadog, OpenTelemetry)
  • ML/AI infrastructure experience is a strong plus
  • Excellent scripting in Python, Go, or Bash
  • Ability to scope and lead infra projects end-to-end
  • Strong incident management and root-cause capability

Responsibilities

  • Own and evolve Nexxa's core infrastructure end-to-end.
  • Design and operate CI/CD pipelines for AI, data, and product teams.
  • Build and maintain infrastructure-as-code for cloud and on-prem/edge deployments.
  • Architect and manage Kubernetes platforms for training, inference, and workloads, including GPU scheduling.
  • Collaborate with data/AI teams on data warehouses, feature stores, and model training infra.
  • Define observability practices: metrics, logging, tracing, alerting across distributed systems.
  • Enforce reliability: SLOs/SLIs, incident response, postmortems, on-call rotations.
  • Design for security/compliance across cloud infra and secrets management.
  • Make cost, latency, reliability tradeoffs, balancing velocity.
  • Help define platform roadmap with engineering leadership.

Skills

Cloud platforms
Kubernetes in prod
Infrastructure-as-code
CI/CD pipelines
Observability stacks
ML/AI infra
Scripting (Python/Go/Bash)
Project leadership
Incident management

Tools

Terraform
Pulumi
GitHub Actions
GitLab CI
CircleCI
Jenkins
ArgoCD

Job description

Nexxa.ai in Canada is seeking a Senior/Staff DevOps Engineer to own and operate the infrastructure for AI and industrial systems at scale. You will oversee compute, networking, storage, and deployment, ensuring production ML workloads run fast, with strong observability and reliability.

You will collaborate with AI, data, and product teams, design CI/CD pipelines, IaC with Terraform/Pulumi, and manage Kubernetes platforms with GPU scheduling.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Staff DevOps Engineer
Staff DevOps Engineer

Nexxa.AI • Toronto

On-site
CAD 130,000 - 190,000
Competitive compensation
Equity package
Senior DevOps Engineer, AI Platform
Senior DevOps Engineer, AI Platform

Triwill Group • Canada

Hybrid
CAD 130,000 - 170,000
Staff AI Platform Engineer — Scale & DX Leader
Staff AI Platform Engineer — Scale & DX Leader

Alygra • Canada

On-site
CAD 150,000 - 190,000
Senior ML Infrastructure Engineer for AI at Scale
Senior ML Infrastructure Engineer for AI at Scale

Jobgether • Toronto

Hybrid
CAD 185,000 - 225,000
Annual bonus
RSU equity
Health benefits
+3
Senior Production AI Systems Engineer
Senior Production AI Systems Engineer

Jaide Health • Toronto

Hybrid
CAD 140,000 - 210,000
Lunch stipend
Health and dental benefits
RRSP matching
+5
AI/ML Infrastructure Engineer
AI/ML Infrastructure Engineer

BULL-IT SOLUTIONS LTD • Montreal

On-site
CAD 100,000 - 130,000
Senior ML Engineer: Scalable AI Infra & Orchestration
Senior ML Engineer: Scalable AI Infra & Orchestration

HelloFresh • Toronto

Hybrid
CAD 170,000 - 190,000
Box discounts
Health & dental benefits
Generous vacation & PTO
+5
Backend AI Engineer
Backend AI Engineer

Nexxa.AI • Toronto

On-site
CAD 120,000 - 190,000
Site Reliability Engineer, AI/ML Infrastructure
Site Reliability Engineer, AI/ML Infrastructure

Boson AI • Toronto

On-site
CAD 100,000 - 130,000
Senior LLMOps Engineer -Cloud / AI Infrastructure
Senior LLMOps Engineer -Cloud / AI Infrastructure

Talent To Hire Inc. • Toronto

On-site
CAD 120,000 - 160,000
Competitive salary
Meaningful equity
Innovative work culture