Senior SRE - Remote, High-Impact Infra

Akka

San Francisco (CA)

On-site

USD 104,000 - 162,000

Full time

14 days+
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Benefits offered by this job

Competitive salary
Health benefits
Professional development
Remote-friendly
Inclusive culture
Work-life balance
Challenging work
Bright minds

Job summary

Akka is seeking a staff-level Site Reliability Engineer to help run and evolve our platform infrastructure. You’ll contribute to architecture decisions for deployment, observability, and security, and help shape team standards.

This is not a watch-keeping role; you will be contributing meaningfully within your first weeks. The core platform uses Kubernetes operators written in Go, with Crossplane, Flux, Terraform, and managed Postgres across cloud regions.

Qualifications

  • 7+ years in infrastructure or platform engineering
  • Experience operating Kubernetes controllers built on controller-runtime in production
  • Terraform and Crossplane in production with Flux and Kustomize
  • Operating managed Postgres in production with point-in-time restore and version upgrades
  • Observability at scale with Prometheus, OpenTelemetry, and long-term metrics stores
  • Securing Kubernetes clusters, service mesh, and secrets management via cloud KMS
  • Production on-call experience and incident ownership
  • Skilled use of LLMs as a tool to sharpen work, not autopilot
  • Strong written communication

Responsibilities

  • Help run and evolve the platform infrastructure
  • Contribute to architecture decisions for deployment, observability, and security
  • Shape team standards and practices, ensuring reliable operations
  • Work hands-on with Kubernetes operators and Go-based controllers
  • Read logs and diagnose root causes using operator status and CRDs
  • Maintain production Postgres across cloud regions with minimal downtime

Skills

Kubernetes
Go
Controller-runtime
Terraform
Crossplane
AWS
Azure
GCP
OpenTelemetry
On-call
LLMs
Communication

Tools

Flux
Kustomize
Postgres

Job description

Akka is seeking a staff-level Site Reliability Engineer to help run and evolve our platform infrastructure. You’ll contribute to architecture decisions for deployment, observability, and security, and help shape team standards.

This is not a watch-keeping role; you will be contributing meaningfully within your first weeks. The core platform uses Kubernetes operators written in Go, with Crossplane, Flux, Terraform, and managed Postgres across cloud regions.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Staff Site Reliability Engineer, Remote Platform
Staff Site Reliability Engineer, Remote Platform

Akka • United States

On-site
USD 150,000 - 190,000
Flexible remote work
Professional development
Competitive salary
+1
Senior SRE Lead: AWS, Kubernetes & Reliability
Senior SRE Lead: AWS, Kubernetes & Reliability

Akoyaexternal • Boston (MA)

On-site
USD 75,768 - 103,320
Senior SRE: Observability, Automation & Scalable Systems
Senior SRE: Observability, Automation & Scalable Systems

Replit • Northern (KY)

Hybrid
USD 140,000 - 190,000
Competitive Salary & Equity
401(k) 4% match (US)
Health, Dental, Vision & Life
+7
Remote Senior SRE Manager - Platform Reliability
Remote Senior SRE Manager - Platform Reliability

Akamai Career Site • United States

On-site
USD 254,000 - 313,000
Healthcare
401K savings plan
Paid time off (PTO)
+4
Remote SRE Engineer - Scale, Automate, Equity Eligible
Remote SRE Engineer - Scale, Automate, Equity Eligible

Akamai Technologies • Cambridge (MA)

Hybrid
USD 75,700 - 136,300
Healthcare
401(k) plan
PTO
+1
Staff SRE: Scale, Observability & Automation Leader
Staff SRE: Scale, Observability & Automation Leader

Replit • Northern (KY)

Hybrid
USD 180,000 - 260,000
Salary & equity
401(k) matching
Health, dental, vision, life
+9
Senior SRE: Scale Infra, Automate, Elevate Reliability
Senior SRE: Scale Infra, Automate, Elevate Reliability

Fathom.ai • United States

On-site
USD 100,000 - 130,000
Competitive compensation
Supportive environment for personal growth
Dynamic and collaborative team
Remote SRE: Automation & Scalable Infrastructure
Remote SRE: Automation & Scalable Infrastructure

Akamai Technologies GmbH • Cambridge (MA)

Hybrid
USD 75,700 - 136,300
Health insurance
401K savings plan
Parental leave
+1
Senior SRE: Platform Reliability & Incident Lead (Remote)
Senior SRE: Platform Reliability & Incident Lead (Remote)

Affirm, Inc. • Town of Poland (NY)

On-site
USD 32,000 - 48,000
Health insurance
Equity rewards
Flexible Spending Wallets
+1
Staff SRE: Scale, Observability & Kubernetes
Staff SRE: Scale, Observability & Kubernetes

Replit • Foster City (CA)

On-site
USD 180,000 - 260,000
Competitive Salary & Equity
401(k) with 4% match
Health, Dental, Vision and Life Ins.
+2