Turn this role into an interview — a resume and cover letter built around what this employer wants.
Kubex is seeking a Technical Lead, Kubernetes Observability to own architecture and development of our telemetry systems across Kubernetes clusters. You will influence how metrics, events, logs and traces are collected, enriched, and used to optimize performance and cost.
The ideal candidate has 7+ years building production software on Kubernetes, deep experience with Prometheus, OpenTelemetry, and Go, and a hands-on leadership style that drives delivery while guiding teams.
Kubex is building the future of autonomous, AI-driven infrastructure optimization. Our platform enables intelligent, policy-driven optimization across Kubernetes, cloud, and GPU-backed environments improving performance, reducing cost, and eliminating waste for some of the world’s most sophisticated technology organizations.
As AI workloads increasingly run on Kubernetes, especially for inference at scale, Kubex is expanding its GPU support to provide more advanced optimization and automation aligned with the unique challenges of GPU-accelerated infrastructure. We combine deep systems expertise, advanced analytics, and patented optimization technology to help customers run AI workloads efficiently and reliably in real-world production environments.
Kubex is seeking a Technical Lead, Kubernetes Observability to own the architecture, technical direction and development of our Kubernetes observability capabilities. This is a senior technical leadership role for someone who can define how telemetry is collected, enriched, processed, and used to support infrastructure optimization across complex customer environments.
You will lead the design and evolution of systems that capture workload behavior, resource usage, performance, and operational context across Kubernetes clusters. This includes making key architectural decisions, establishing engineering patterns, prototyping new approaches, and contributing directly to the most critical and technically challenging parts of the implementation.
This role combines high-level ownership with strong hands-on engineering. You will use modern AI-assisted development workflows to accelerate implementation, testing, investigation, and iteration, while remaining accountable for system design, technical decisions, code quality, and production reliability.
The ideal candidate has deep experience building observability or telemetry systems for Kubernetes and is comfortable working across metrics, events, logs, traces, workload metadata, and distributed data collection. Experience with GPU or AI workloads is valuable but not required. You will help build the observability foundation that supports Kubex’s current Kubernetes optimization capabilities and its continued expansion into GPU and AI infrastructure.