Staff Software Engineer - Kubernetes Observability

Kubex

Canada

On-site

CAD 140,000 - 190,000

Full time

10 days ago
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Benefits offered by this job

Remote-friendly culture
Competitive compensation
Equity
Benefits package

Job summary

Kubex is seeking a Technical Lead, Kubernetes Observability to own architecture and development of our telemetry systems across Kubernetes clusters. You will influence how metrics, events, logs and traces are collected, enriched, and used to optimize performance and cost.

The ideal candidate has 7+ years building production software on Kubernetes, deep experience with Prometheus, OpenTelemetry, and Go, and a hands-on leadership style that drives delivery while guiding teams.

Qualifications

  • 7+ years of software engineering experience with Kubernetes.
  • Experience designing and building observability/telemetry systems.
  • Strong ownership of architecture while remaining hands-on.

Responsibilities

  • Design telemetry collection and processing for Kubernetes.
  • Lead hands-on implementation and code contributions.
  • Collaborate with senior engineers and product managers.
  • Prototype and productionize AI/workload observability approaches.
  • Identify opportunities to extend observability to GPUs/AI workloads.

Skills

Kubernetes observability
Prometheus
OpenTelemetry
Go
Distributed systems
Technical leadership

Tools

Prometheus
OpenTelemetry

Job description

Kubex is building the future of autonomous, AI-driven infrastructure optimization. Our platform enables intelligent, policy-driven optimization across Kubernetes, cloud, and GPU-backed environments improving performance, reducing cost, and eliminating waste for some of the world’s most sophisticated technology organizations.

As AI workloads increasingly run on Kubernetes, especially for inference at scale, Kubex is expanding its GPU support to provide more advanced optimization and automation aligned with the unique challenges of GPU-accelerated infrastructure. We combine deep systems expertise, advanced analytics, and patented optimization technology to help customers run AI workloads efficiently and reliably in real-world production environments.

Role Overview

Kubex is seeking a Technical Lead, Kubernetes Observability to own the architecture, technical direction and development of our Kubernetes observability capabilities. This is a senior technical leadership role for someone who can define how telemetry is collected, enriched, processed, and used to support infrastructure optimization across complex customer environments.

You will lead the design and evolution of systems that capture workload behavior, resource usage, performance, and operational context across Kubernetes clusters. This includes making key architectural decisions, establishing engineering patterns, prototyping new approaches, and contributing directly to the most critical and technically challenging parts of the implementation.

This role combines high-level ownership with strong hands-on engineering. You will use modern AI-assisted development workflows to accelerate implementation, testing, investigation, and iteration, while remaining accountable for system design, technical decisions, code quality, and production reliability.

The ideal candidate has deep experience building observability or telemetry systems for Kubernetes and is comfortable working across metrics, events, logs, traces, workload metadata, and distributed data collection. Experience with GPU or AI workloads is valuable but not required. You will help build the observability foundation that supports Kubex’s current Kubernetes optimization capabilities and its continued expansion into GPU and AI infrastructure.

Key Responsibilities
  • Lead the design of systems that collect telemetry and improve performance of inference workloads.
  • Contribute directly to production code, remaining deeply hands-on in the design, implementation, and evolution of core platform components.
  • Collaborate closely with other senior engineers, product managers and engineering leadership to coordinate and execute complex software development initiatives.
  • Prototype, validate, and productionize new technical approaches related to AI workload observability and performance optimization.
  • Identify opportunities to extend Kubex’s value beyond inference workloads, including potential future optimizations for training or hybrid workloads.
Required Qualifications
  • 7+ years of professional software engineering experience, including significant experience building production software on Kubernetes.
  • Strong experience designing or building observability and telemetry solutions using technologies such as Prometheus, OpenTelemetry, or similar platforms.
  • Deep understanding of Kubernetes workloads, resource management, API interactions, and the operational challenges of running software across diverse customer clusters.
  • Strong coding skills, preferably in Go, with experience building scalable, reliable, and testable distributed systems.
  • Demonstrated ability to own technical architecture while remaining hands-on with implementation, prototyping, debugging, and production delivery.
Preferred Qualifications
  • Experience building collectors, exporters, agents, or telemetry-processing pipelines for Kubernetes environments.
  • Knowledge of performance analysis, resource optimization, scheduling, or infrastructure efficiency.
  • Exposure to GPU-backed infrastructure, AI inference workloads, or GPU observability technologies.
Why Join Kubex?
  • Play a key role in shaping the future of AI infrastructure optimization.
  • Work on technically challenging problems at the intersection of Kubernetes, GPUs, and AI workloads.
  • Collaborate with a highly experienced, deeply technical team.
  • Influence product direction, architecture, and external technical positioning.
  • Flexible, remote-first culture focused on impact and innovation.
  • Competitive compensation, equity, and benefits.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Kubernetes Observability Tech Lead
Kubernetes Observability Tech Lead

Kubex • Canada

On-site
CAD 140,000 - 190,000
Remote-friendly culture
Competitive compensation
Equity
+1
[8SN] Senior Site Reliability Engineer (SRE) – Kubernetes
[8SN] Senior Site Reliability Engineer (SRE) – Kubernetes

Worky • Montreal (administrative region)

On-site
CAD 120,000 - 170,000
Laptop
Flexible work arrangements
Professional development and training
Sr. Solutions Architect (Post-sales)
Sr. Solutions Architect (Post-sales)

Rafay • Toronto

On-site
CAD 110,000 - 150,000
Competitive salary
Robust benefits
Attractive stock options
+1
Senior Software Developer
Senior Software Developer

JSI • Eastern Ontario

Hybrid
USD 64,000 - 100,000
Senior DevOps Engineer
Senior DevOps Engineer

MarkiTech • Toronto

On-site
CAD 120,000 - 180,000
Staff DevOps Engineer
Staff DevOps Engineer

Engg • Toronto

On-site
CAD 140,000 - 210,000
Competitive compensation package
Equity or stock options
System Engineer - Golang & Kubernetes CRDs
System Engineer - Golang & Kubernetes CRDs

PortBlueSky • British Columbia

On-site
CAD 90,000 - 120,000
Remote-first flexibility
Deep systems engineering work
Strong technical environment
+2
AI/ML Engineer
AI/ML Engineer

Katalyst Data Management • Calgary

Hybrid
CAD 120,000 - 180,000
Hybrid schedule
Staff Data Engineer
Staff Data Engineer

Katalyze AI, Inc. • Toronto

On-site
CAD 120,000 - 180,000
Senior DevOps Engineer
Senior DevOps Engineer

MarkiTech.AI • Toronto

On-site
CAD 140,000 - 180,000