Remote Cloud SRE & NOC Reliability Lead

AXON Networks

Irvine (CA)

On-site

USD 160,000 - 200,000

Full time

2 days ago
Be an early applicant
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Job summary

AXON Networks is seeking a Site Reliability Engineer to improve the availability, performance, scalability and recoverability of our cloud solutions. You will blend software engineering with hands-on NOC operations to make the cloud-to-device service path observable, supportable and resilient at fleet scale.

Join a team that partners with Support, Operations, cloud and DevOps to reduce on-call toil, develop production-grade tooling, and drive reliability improvements across APIs, Kubernetes,

Qualifications

  • 5+ years of experience in site reliability engineering, production engineering, DevOps, cloud infrastructure, systems engineering or a closely related role.
  • Strong software or automation skills in Python, Go, Java, Bash or a comparable language, with experience producing maintainable operational code.
  • Hands-on experience operating distributed production systems in a public cloud environment and troubleshooting across application, infrastructure, network and device-integration layers.
  • Experience with Google Cloud Platform, Oracle Cloud Infrastructure and production Kubernetes environments.
  • Experience with infrastructure as code and delivery tooling such as Terraform, Helm, Git-based CI/CD and policy-as-code.
  • Strong Linux, containers and Kubernetes fundamentals, including deployment behavior, resource management, networking and failure diagnosis.
  • Strong troubleshooting & debugging skills in Kubernetes platforms.
  • Experience with modern observability practices and tools across metrics, logs, traces, alerting, dashboards and synthetic monitoring.
  • Familiarity with Prometheus, Grafana, OpenTelemetry or equivalent observability ecosystems.
  • Familiarity with Apache Pulsar or similar distributed messaging and streaming platforms handling requests from millions of devices.
  • Experience participating in an on-call rotation and responding effectively to high-severity, customer-impacting production incidents.
  • Working knowledge of SLOs, error budgets, capacity planning, resilience engineering, change safety and blameless incident learning.
  • Strong networking knowledge, including TCP/IP, DNS, DHCP, TLS, routing, NAT, load balancing and systematic packet- or session-level troubleshooting.
  • Clear communication, disciplined documentation and the ability to collaborate across NOC, cloud, DevOps, firmware and service-provider teams.
  • Bachelor’s degree in computer science, engineering or equivalent practical experience.

Responsibilities

  • Own reliability outcomes for assigned cloud services.
  • Improve observability, capacity, resilience and recovery.
  • Define and operationalize service-level indicators, service-level objectives and actionable alerting.
  • Automate repetitive NOC work and create safe, testable mechanisms for diagnosis, recovery, device operations and routine production changes.
  • Lead technically during incidents, drive evidence-based learning and ensure high-value corrective actions completed.

Skills

SRE/DevOps
Programming
Kubernetes
Linux fundamentals
Observability
Networking
On-call experience

Education

Bachelor’s degree

Tools

Prometheus
Grafana
OpenTelemetry
Terraform
Helm
Git CI/CD
Apache Pulsar

Job description

AXON Networks is seeking a Site Reliability Engineer to improve the availability, performance, scalability and recoverability of our cloud solutions. You will blend software engineering with hands-on NOC operations to make the cloud-to-device service path observable, supportable and resilient at fleet scale.

Join a team that partners with Support, Operations, cloud and DevOps to reduce on-call toil, develop production-grade tooling, and drive reliability improvements across APIs, Kubernetes,

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Site Reliable Engineer (US - Remote)
Site Reliable Engineer (US - Remote)

AXON Networks • Irvine (CA)

On-site
USD 160,000 - 200,000
Remote Network Operations Lead – Automation & Reliability
Remote Network Operations Lead – Automation & Reliability

AXON Networks • Irvine (CA)

On-site
USD 110,000 - 160,000
Remote within North America
Remote OSS Platform Engineer — Automation & NOC Tools
Remote OSS Platform Engineer — Automation & NOC Tools

AXON-Networks • Irvine (CA)

Remote
USD 90,000 - 150,000
Network Operations & Automation Lead
Network Operations & Automation Lead

AXON-Networks • Irvine (CA)

Remote
USD 110,000 - 160,000
OSS Platform Engineer - Automation & NOC Tools
OSS Platform Engineer - Automation & NOC Tools

AXON Networks • Irvine (CA)

On-site
USD 90,000 - 150,000
Fully remote (North America)
Remote Senior Network Reliability Engineer (SRE)
Remote Senior Network Reliability Engineer (SRE)

Gainbridge • Zionsville (IN), Northern (KY)

On-site
USD 135,000 - 190,000
Health Insurance
Dental Insurance
Vision Insurance
+4
Senior SRE: Platform Reliability & Observability Lead
Senior SRE: Platform Reliability & Observability Lead

Prepared • New York (NY)

Hybrid
USD 141,000 - 217,000
Competitive salary and 401k with match
Cloud-Native DevOps Engineer: Scale, Automation & CI/CD
Cloud-Native DevOps Engineer: Scale, Automation & CI/CD

AXON Networks • Irvine (CA)

On-site
USD 100,000 - 140,000
SRE II - Cloud Automation & Platform Engineer
SRE II - Cloud Automation & Platform Engineer

Koitecc Solutions • Boston (MA), Northern (KY)

Hybrid
USD 116,000 - 165,000
Competitive salary
401k with employer match
Discretionary paid time off
+7
Senior Site Reliability Lead - Cloud & Automation
Senior Site Reliability Lead - Cloud & Automation

NetApp • Morrisville (NC)

Hybrid
USD 170,000 - 253,000
Health Insurance
Life Insurance
Retirement or Pension Plans
+5