Senior SRE: Platform Reliability & Security Lead

WellSaid Labs

United States

On-site

USD 120,000 - 180,000

Full time

14 days+
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Job summary

WellSaid Labs is seeking a Site Reliability Engineer to run and scale our production platform, focusing on availability and performance of core services and data pipelines. You will build and operate monitoring, tracing and alerting, while driving security initiatives and responding to incidents with thorough root cause analysis.

A strong emphasis on automated deployments and Kubernetes-based environments will be central to daily work.

Qualifications

  • Experience with cloud environments (AWS or GCP) and container tech.
  • 5+ years with a modern programming language (Go, TypeScript, Python, etc.).
  • Experience with monitoring tools such as Grafana and Prometheus.
  • Experience with GitOps/continuous delivery tooling (ArgoCD, Spacelift, Terraform Cloud).
  • Strong understanding of Infrastructure as Code (Terraform, Pulumi, etc.).
  • Ability to debug and solve issues in complex production environments.
  • Experience building and troubleshooting Kubernetes environments.
  • Fluency with UNIX shell to analyze logs and perform operational tasks.
  • Understanding of profiling applications and databases for performance issues.

Responsibilities

  • Run the production platform and large-scale projects to improve availability and performance.
  • Build and run monitoring, tracing and alerting infrastructure.
  • Drive platform security initiatives and preventative reliability improvements.
  • Lead incident response, root cause analysis and recovery efforts.
  • Operate with high availability and redundancy as a core principle.
  • Respond to alerts and participate in on-call rotations.
  • Improve deployment processes for fast, safe and simple code changes.
  • Deliver a stable, scalable product platform enabling fast engineering delivery.
  • Identify novel load-handling approaches for resource-intensive apps.

Skills

Cloud environments
Kubernetes
Docker
Go / Typescript / Python
Grafana / Prometheus
GitOps / CD tooling
Terraform / Pulumi / IaC
Terraform Cloud
ArgoCD / Spacelift
Unix shell
Performance profiling

Tools

AWS
GCP
Kubernetes
Docker
Grafana
Prometheus
ArgoCD
Spacelift
Terraform
Pulumi
Terraform Cloud
Terraform / IaC tooling
Unix shell

Job description

WellSaid Labs is seeking a Site Reliability Engineer to run and scale our production platform, focusing on availability and performance of core services and data pipelines. You will build and operate monitoring, tracing and alerting, while driving security initiatives and responding to incidents with thorough root cause analysis.

A strong emphasis on automated deployments and Kubernetes-based environments will be central to daily work.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Remote Senior SRE: Reliability, Security & Scale
Remote Senior SRE: Reliability, Security & Scale

WellSaid Labs, Inc. • Northern (KY)

Hybrid
USD 140,000 - 190,000
Competitive salary
Stock options
Medical, dental, and vision insurance
+4
Senior SRE: Scale Reliability, Observability & Resilience
Senior SRE: Scale Reliability, Observability & Resilience

Early Warning Services LLC • Scottsdale (AZ)

Hybrid
USD 106,000 - 130,000
Healthcare Coverage
401(k) Retirement Plan
Paid Time Off
+2
Senior SRE: Build Scalable, Reliable Platforms — Remote
Senior SRE: Build Scalable, Reliable Platforms — Remote

Nord Security • Town of Poland (NY)

On-site
USD 140,000 - 200,000
Premium healthcare
Work from anywhere
Mentorship programs
+3
Senior SRE: Incident & Kubernetes Reliability Lead
Senior SRE: Incident & Kubernetes Reliability Lead

Unique System Skills • United States

Remote
USD 140,000 - 190,000
Senior SRE - Hybrid, Platform Reliability Lead
Senior SRE - Hybrid, Platform Reliability Lead

TransUnion LLC • Reston (VA)

Hybrid
USD 112,500 - 187,500
Day-one medical, dental, vision
Company-paid basic life/AD&D
12 weeks paid parental leave
+2
Senior SRE: Build Resilient, Scalable Systems
Senior SRE: Build Resilient, Scalable Systems

Methodic • San Francisco (CA)

On-site
USD 140,000 - 210,000
Senior SRE: Scalable Infra, Observability & Automation
Senior SRE: Scalable Infra, Observability & Automation

Early Warning • Scottsdale (AZ)

Hybrid
USD 106,000 - 156,000
Healthcare Coverage
401(k) Company Match
Paid Time Off
+2
Senior Engineer - Site Reliability Engineering
Senior Engineer - Site Reliability Engineering

LSEG (London Stock Exchange Group) • Allen (TX)

On-site
USD 140,000 - 190,000
Senior SRE Lead - Platform Resiliency & Observability
Senior SRE Lead - Platform Resiliency & Observability

SIMARN Solutions • Charlotte (NC)

On-site
USD 120,000 - 160,000
Senior SRE: Platform Reliability & Incident Lead (Remote)
Senior SRE: Platform Reliability & Incident Lead (Remote)

Affirm, Inc. • Town of Poland (NY)

On-site
USD 32,000 - 48,000
Health insurance
Equity rewards
Flexible Spending Wallets
+1