Senior SRE: GitOps & Kubernetes Reliability (Remote)

Akuity

United States

Remote

USD 120,000 - 160,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Competitive compensation
Equity participation
Flexible time off
Home office stipend
Work from anywhere in US time zones

Job summary

Akuity is looking for a Senior SRE to ensure the reliability of its platform. This high-ownership role involves collaborating with engineering and infrastructure teams to build resilience into our systems.

With a focus on observability, incident response, and continuous improvement of SLI/SLO/SLA definitions, the ideal candidate will have significant experience in production operations, particularly within a SaaS environment. The position is fully remote, open to candidates in US time zones, and offers competitive compensation and equity participation.

Qualifications

  • 5+ years of SRE, platform engineering, or production operations experience in a SaaS environment.
  • Deep hands-on Kubernetes expertise.
  • Strong AWS fundamentals across compute, networking, and storage.
  • Experience defining and operating against SLOs in production.
  • Proficiency with observability tooling.
  • Solid scripting and automation skills in Go, Python, or Bash.
  • Strong written communication.

Responsibilities

  • Own SLI/SLO/SLA definitions for the Akuity SaaS platform.
  • Design and maintain observability systems across AWS infrastructure.
  • Identify reliability gaps and lead post-mortems.
  • Participate in on-call rotation and act as incident commander.
  • Build and maintain runbooks and incident playbooks.

Skills

Kubernetes expertise
AWS fundamentals
Observability tooling
Scripting and automation skills
Strong written communication

Tools

Prometheus
Grafana
OpenTelemetry
Terraform

Job description

Akuity is looking for a Senior SRE to ensure the reliability of its platform. This high-ownership role involves collaborating with engineering and infrastructure teams to build resilience into our systems.

With a focus on observability, incident response, and continuous improvement of SLI/SLO/SLA definitions, the ideal candidate will have significant experience in production operations, particularly within a SaaS environment. The position is fully remote, open to candidates in US time zones, and offers competitive compensation and equity participation.

Get your free, confidential resume review.
or drag and drop your file here.