Platform SRE - Scale Kubernetes, Cloud Infra & Reliability

Specter

San Francisco (CA)

On-site

USD 180,000 - 230,000

Full time

11 days ago

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Specter is hiring a Platform Site Reliability Engineer to own the health, reliability, and scalability of the cloud platform behind our connected sensor fleet. You will operate Kubernetes-based infrastructure, manage cloud resources with Terraform, strengthen observability, and support safe deployments across teams.

You’ll collaborate with application, AI, embedded systems, and fleet teams on AWS and Kubernetes.

Qualifications

  • Strong Linux systems knowledge and experience diagnosing production systems.
  • Hands-on experience operating Kubernetes in production, including networking, storage, resource management, upgrades, and troubleshooting.
  • Strong experience using Terraform to manage production cloud infrastructure.
  • Experience with AWS, including IAM, networking, compute, storage, and EKS.
  • Solid networking fundamentals, including DNS, load balancing, firewalls, VPNs, subnets, and routing.

Responsibilities

  • Reactive - Triage & Recovery: Debug production issues across Kubernetes clusters, Linux systems, AWS, networking, and workloads; Lead incidents from detection through recovery with cross-team coordination; Participate in on-call rotations and ensure durable fixes.
  • Systems Builder - Close the Loop: Build, operate, and improve our Kubernetes platform and AWS infrastructure; Manage production infrastructure with Terraform, including reusable modules, automated validation, and safe change workflows; Reduce toil through automation and improved deployment tooling, CI/CD, and developer workflows.
  • Observability Owner - Platform Visibility: Design and improve observability with logging, metrics, tracing, dashboards, and alerting; Define service-level indicators and close telemetry gaps; Develop runbooks and incident-response procedures.

Skills

Linux
Kubernetes
Terraform
AWS
Networking
Scripting
Incident management

Tools

Terraform

Job description

Specter is hiring a Platform Site Reliability Engineer to own the health, reliability, and scalability of the cloud platform behind our connected sensor fleet. You will operate Kubernetes-based infrastructure, manage cloud resources with Terraform, strengthen observability, and support safe deployments across teams.

You’ll collaborate with application, AI, embedded systems, and fleet teams on AWS and Kubernetes.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Cloud Platform SRE Engineer | Remote, Kubernetes & IaC
Cloud Platform SRE Engineer | Remote, Kubernetes & IaC

Spectraforce Technologies • Raleigh (NC)

On-site
USD 96,000 - 165,000
Platform Site Reliability Engineer
Platform Site Reliability Engineer

Specter • San Francisco (CA)

On-site
USD 180,000 - 230,000
Senior Platform Engineer: Cloud, CI/CD & Kubernetes
Senior Platform Engineer: Cloud, CI/CD & Kubernetes

Spectraforce Technologies • SeaTac (WA)

On-site
USD 140,000 - 190,000
Edge & Cloud SRE: Fleet Reliability & Observability
Edge & Cloud SRE: Fleet Reliability & Observability

Specter • San Francisco (CA)

On-site
USD 120,000 - 160,000
Kubernetes Platform SRE – Cloud Infra & Automation
Kubernetes Platform SRE – Cloud Infra & Automation

Okta • Bellevue (WA)

On-site
USD 174,000 - 214,000
Kubernetes Platform SRE - Cloud-Native Infra Lead
Kubernetes Platform SRE - Cloud-Native Infra Lead

United States Digital Space LLC • Washington

Hybrid
USD 174,000 - 214,000
Health, dental & vision
401(k) plan
Paid time off
Platform SRE Engineer — Kubernetes, IaC & Cloud
Platform SRE Engineer — Kubernetes, IaC & Cloud

Menlo Security • United States

On-site
USD 80,000 - 120,000
Senior Site Reliability Engineer - Kubernetes, AWS & GitOps
Senior Site Reliability Engineer - Kubernetes, AWS & GitOps

Pinterest • United States

On-site
USD 139,000 - 288,000
Senior Cloud SRE: Reliability, Automation & Platform
Senior Cloud SRE: Reliability, Automation & Platform

Carrier • Atlanta (GA)

On-site
USD 96,000 - 192,000
Health Care Benefits
Retirement Benefits
Time off and Leave: Paid vacation days
+1
Senior SRE: Kubernetes Platform Reliability (Remote)
Senior SRE: Kubernetes Platform Reliability (Remote)

Circle Internet Management Services LLC • California (MO)

On-site
USD 153,000 - 205,000