Senior DevOps Engineer, AI Platform

Webhosting

Canada

Hybrid

CAD 120,000 - 160,000

Full time

22 hours ago
Be an early applicant
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Job summary

Network Solutions is seeking a hands-on Senior DevOps Engineer to build and operate the infrastructure powering our AI platforms, agent runtimes, and web applications. You will work across cloud environments to ensure scalable, secure production systems.

You’ll translate designs into production infrastructure in Azure and Oracle Cloud, implement end-to-end observability, and own incident response, cost optimization, and reliability.

Qualifications

  • Experience building production cloud infrastructure.
  • Strong Kubernetes and cloud ops experience.
  • Ability to design scalable, observable platforms.

Responsibilities

  • Translate application and platform designs into production-ready cloud infrastructure with minimal supervision.
  • Design, provision, operate, and troubleshoot Kubernetes environments, primarily AKS and Oracle Kubernetes Engine.
  • Support AI workloads including LiteLLM gateways, Python agent runtimes, RAG workers, MCP services, background workers, and asynchronous processing pipelines.
  • Design and manage ingress and egress networking, load balancers, DNS, TLS, private connectivity, routing, NAT, firewalls, network policies, and service-to-service communication.
  • Build and operate infrastructure for web applications and backend services, including APIs, databases, caches, queues, scheduled jobs, and event-driven workloads.
  • Build and maintain CI/CD pipelines using Jenkins and Bitbucket, integrating Docker, Helm, Kubernetes, ArgoCD, and container registries.
  • Automate infrastructure provisioning and configuration using Terraform, Helm, Kubernetes manifests, Python, Bash, and related tooling.
  • Implement end-to-end observability using metrics, logs, distributed tracing, dashboards, alerts, health checks, and SLOs.
  • Own production readiness, incident troubleshooting, root cause analysis, scalability, reliability, and infrastructure cost optimization.
  • Create reusable infrastructure patterns that allow engineering teams to launch new services quickly and consistently.
  • You do not need to be a full time application developer, but you should understand how modern backend systems work and be able to troubleshoot across application and infrastructure boundaries.

Skills

Kubernetes
CI/CD
Cloud infra
Python
Bash
Terraform
Helm
ArgoCD
Docker
Networking
Observability
SRE

Tools

Jenkins
Bitbucket
Docker
Kubernetes
Terraform
Helm
ArgoCD
Azure
OCI

Job description

At Network Solutions, we’ve been trusted for decades to help people get online and stay ahead. We’ve been here since the beginning of the internet, and we’re still building for what comes next.

As the original digital identity authority, we help secure domain names, protect brands, and safeguard the infrastructure businesses rely on. We empower our customers to own and manage the assets that define them online, while delivering enterprise-grade security to protect against virtual threats. Our team leverages modern, AI-accelerated tools to streamline how businesses manage their digital presence, making the most of our decades of experience.

The Network Solutions team is here to help online businesses protect what’s theirs and build for tomorrow. That’s why millions trust us to protect their domains, brands, and websites every day.

The impact you’ll make

We are looking for a hands-on Senior DevOps Engineer to build and operate the infrastructure powering our AI platforms, agent runtimes, web applications, backend services, APIs, and shared platform capabilities. Our environment includes a centralized LLM gateway, Python-based agent runtimes, RAG workers, MCP services, asynchronous processing, databases, caches, queues, and observability services.

You will work with AI engineers, application engineers, and architects who define technical designs, then independently translate those designs into reliable, scalable, secure, and observable production infrastructure across Microsoft Azure and Oracle Cloud Infrastructure.

What you’ll do
  • Translate application and platform technical designs into production ready cloud infrastructure with minimal supervision.
  • Design, provision, operate, and troubleshoot Kubernetes environments, primarily Azure Kubernetes Service and Oracle Kubernetes Engine.
  • Support AI workloads including LiteLLM based gateways, Python agent runtimes, RAG workers, MCP services, background workers, and asynchronous processing pipelines.
  • Design and manage ingress and egress networking, load balancers, DNS, TLS, private connectivity, routing, NAT, firewalls, network policies, and service to service communication.
  • Build and operate infrastructure for web applications and backend services, including APIs, databases, caches, queues, scheduled jobs, and event driven workloads.
  • Build and maintain CI/CD pipelines using Jenkins and Bitbucket, integrating Docker, Helm, Kubernetes, ArgoCD, and container registries.
  • Automate infrastructure provisioning and configuration using Terraform, Helm, Kubernetes manifests, Python, Bash, and related tooling.
  • Implement end to end observability using metrics, logs, distributed tracing, dashboards, alerts, health checks, and SLOs.
  • Own production readiness, incident troubleshooting, root cause analysis, scalability, reliability, and infrastructure cost optimization.
  • Create reusable infrastructure patterns that allow engineering teams to launch new services quickly and consistently.
  • You do not need to be a full time application developer, but you should understand how modern backend systems work and be able to troubleshoot across application and infrastructure boundaries.
How you’ll work

You will frequently receive a technical design for a new AI workload, application, backend service, or platform capability. From that design, you should be able to independently determine and implement the infrastructure needed to run it in production.

What will make you stand out

AI or machine learning infrastructure experience is helpful, but not required. Strong experience with Kubernetes, web and backend infrastructure, networking, queues, databases, CI/CD, observability, and production cloud operations is the foundation for this role.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior DevOps Engineer, AI Platform
Senior DevOps Engineer, AI Platform

Triwill Group • Canada

Hybrid
CAD 130,000 - 170,000
Staff Engineer (AI & Engineering)
Staff Engineer (AI & Engineering)

EQ Bank • Toronto

On-site
CAD 140,000 - 190,000
AI Engineer
AI Engineer

Valsoft Corporation • Canada

On-site
CAD 100,000 - 140,000
DevOps Solution Architect
DevOps Solution Architect

Randstad Digital Americas • Mississauga

On-site
CAD 120,000 - 180,000
Staff Engineer (AI & Engineering)
Staff Engineer (AI & Engineering)

EQ Bank | Canada's Challenger Bank • Toronto

On-site
CAD 150,000 - 210,000
Staff Engineer (AI & Engineering)
Staff Engineer (AI & Engineering)

Kinvie • Toronto

On-site
CAD 140,000 - 190,000
Senior AI Platform Operations Engineer
Senior AI Platform Operations Engineer

EQ Bank • Toronto

On-site
CAD 120,000 - 160,000
Forward Deployed AI Engineer
Forward Deployed AI Engineer

EQ Bank • Toronto

On-site
CAD 140,000 - 190,000
Senior Azure DevOps Engineer
Senior Azure DevOps Engineer

High Tech Genesis Inc. • Canada

Hybrid
CAD 120,000 - 180,000
Staff Software Engineer
Staff Software Engineer

kadence • Toronto

Hybrid
CAD 150,000 - 210,000
Health insurance