Senior Observability Platform SRE

Koitecc Solutions

Seattle, Northern (WA, KY)

Hybrid

USD 134,000 - 215,000

Full time

14 days+
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Benefits offered by this job

Competitive salary & 401k
Discretionary time off
Parental leave

Job summary

Axon is seeking a Site Reliability Engineer to join the Observability team in Seattle. You will own distributed tracing, logging, and metrics infrastructure across multi-environment deployments, collaborating with engineering teams to raise reliability and developer experience.

The role emphasizes hands-on work with OpenTelemetry, Grafana, Loki, Cortex, and Terraform, plus GitOps workflows and on-call readiness. A strong Linux background is essential.

Qualifications

  • Bachelor's Degree in Computer Science, Engineering, or an equivalent highly technical field.
  • 7+ years of experience in SRE, platform engineering, or infrastructure engineering.
  • Experience with agentic AI tooling or building LLM-powered developer tools.
  • Strong Linux systems fundamentals and comfort working in Kubernetes-based environments.
  • Hands-on experience with one or more components of the LGTM stack: Loki, Grafana, Tempo/Jaeger, or Mimir/Cortex.
  • Experience with infrastructure as code - Terraform strongly preferred, CDK is a plus.
  • Experience with any of: Golang, Python, or Java.
  • United States citizen - able to gain CJIS clearance for full US production access.

Responsibilities

  • Own and evolve Axon's distributed tracing infrastructure, including Jaeger and OpenTelemetry-based instrumentation, driving adoption across Axon's service-oriented architecture
  • Build and operate Axon's log aggregation platform (Grafana Loki + Alloy), expanding use cases beyond Kubernetes event logs and reducing organizational dependency on expensive third-party log tooling (including Splunk)
  • Maintain and improve Axon's metrics infrastructure (Cortex, Prometheus, Grafana) - the foundation for alerting, dashboards, and SLO tracking across all of Axon's environments
  • Write internal tooling and automation that makes observability self-service: toolkit commands, agentic on-call helpers, runbook generation, and dashboard scaffolding
  • Manage observability infrastructure as code via Terraform, CDK, ArgoCD, and Helm - including capacity management, cybersecurity requirements and compliance, and on-call rotation participation
  • Work directly with engineering teams across Axon to define instrumentation standards, drive tracing adoption, and help teams build meaningful SLOs for their services

Skills

Linux fundamentals
Kubernetes
Observability concepts
Infrastructure as code
Golang
Python
Java
CJIS clearance

Education

Bachelor's Degree in Computer Science or related field

Tools

Terraform
CDK
ArgoCD
Helm
Grafana
Loki
Tempo/Jaeger
Cortex/Mimir
OpenTelemetry

Job description

Axon is seeking a Site Reliability Engineer to join the Observability team in Seattle. You will own distributed tracing, logging, and metrics infrastructure across multi-environment deployments, collaborating with engineering teams to raise reliability and developer experience.

The role emphasizes hands-on work with OpenTelemetry, Grafana, Loki, Cortex, and Terraform, plus GitOps workflows and on-call readiness. A strong Linux background is essential.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior SRE – Observability Platform & Distributed Systems
Senior SRE – Observability Platform & Distributed Systems

Koitecc Solutions • Boston (MA)

Hybrid
USD 134,250 - 214,800
Competitive salary
401k with employer match
Discretionary time off
+5
Senior SRE: Platform Reliability & Observability Lead
Senior SRE: Platform Reliability & Observability Lead

Prepared • New York (NY)

Hybrid
USD 141,000 - 217,000
Competitive salary and 401k with match
Site Reliability Engineer
Site Reliability Engineer

BlueSky Resource Solutions • Duluth (GA)

On-site
USD 120,000 - 180,000
SRE II: CloudOps Automation & Production Reliability
SRE II: CloudOps Automation & Production Reliability

Out in Science, Technology, Engineering, and Mathematics • Washington

Hybrid
USD 115,000 - 165,000
Competitive salary
401k with employer match
Discretionary paid time off
+7
Senior SRE: Observability, Automation & Scalable Systems
Senior SRE: Observability, Automation & Scalable Systems

Replit • Northern (KY)

Hybrid
USD 140,000 - 190,000
Competitive Salary & Equity
401(k) 4% match (US)
Health, Dental, Vision & Life
+7
Observability Engineer / Site Reliability Engineer
Observability Engineer / Site Reliability Engineer

Ontrac Solutions • New York (NY)

On-site
USD 120,000 - 190,000
Remote Cloud SRE & NOC Reliability Lead
Remote Cloud SRE & NOC Reliability Lead

AXON Networks • Irvine (CA)

On-site
USD 160,000 - 200,000
Senior SRE: Kubernetes Platform Engineer (Hybrid)
Senior SRE: Kubernetes Platform Engineer (Hybrid)

Axon • Boston (MA)

Hybrid
USD 134,000 - 215,000
401k with employer match
Discretionary time off
Parental leave
+4
Senior SRE: Observability & Incident Champion (Azure/AWS)
Senior SRE: Observability & Incident Champion (Azure/AWS)

Hidden Road • New York (NY)

Hybrid
USD 160,000 - 200,000
Senior SRE - Remote, High-Impact Infra
Senior SRE - Remote, High-Impact Infra

Akka • San Francisco (CA)

On-site
USD 104,000 - 162,000
Competitive salary
Health benefits
Professional development
+5