Senior SRE, Platform Software Engineer

Jobtailor

California (MO)

On-site

USD 150,000 - 200,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Jobtailor is seeking a senior software engineer to translate architect-approved designs into production code that ships through GitOps and the CI/CD pipeline. You will build and operate bounded contexts of the NeoCloud SRE platform.

You will own 1-2 services including collection, alert, correlation, and SLO frameworks, and implement fault-prediction and remediation workflows. You will also manage hardware lifecycle and data center operations, while developing identity, secrets, tenant

Qualifications

  • 7+ years of production software engineering experience, including on-call work.
  • Proficiency in at least one systems language: Go, Rust, or Java.
  • Python proficiency for tooling and SDK work.
  • Strong grasp of at-least-once vs. exactly-once trade-offs, idempotency, back-pressure, leader election, consistent hashing, gossip, and fan-out.
  • Experience at production scale with observability tools such as Prometheus/OpenTelemetry/Jaeger or Loki/Tempo.
  • Hands-on experience with Argo, Flux, Helm, Kustomize, Cosign signing, and signed-bundle promotions.
  • Proven experience writing controllers/CRDs with watch-cache, leader election, and reconcile loops.
  • Experience executing end-to-end mTLS bootstrap with certificate rotation.
  • Hands-on experience with Vault or cloud KMS (AWS/GCP).
  • Ability to read Prometheus query plans, write recording rules, and join per-tenant telemetry with analytics data.

Responsibilities

  • Take an architect-approved design and turn it into production code that ships through GitOps + the CICD release pipeline.
  • Build and operate one or more bounded contexts of the NeoCloud SRE platform.
  • Own 1-2 of various services including collection, alert, correlation, and SLO frameworks.
  • Implement fault-prediction and remediation workflows.
  • Manage hardware lifecycle and data center operations.
  • Develop services for identity, secrets, tenant configuration, and customer bridging.

Skills

Go (preferred)
Rust
Java
Python
On-call experience

Tools

Prometheus
VictoriaMetrics
Mimir
Thanos
Loki
Elasticsearch
OpenTelemetry
Argo
Flux
Helm
Kustomize
Cosign signing
HashiCorp Vault
AWS KMS
GCP KMS

Job description

Responsibilities
  • Take an architect-approved design and turn it into production code that ships through GitOps + the CICD release pipeline.
  • Build and operate one or more bounded contexts of the NeoCloud SRE platform.
  • Own 1-2 of various services including collection, alert, correlation, and SLO frameworks.
  • Implement fault-prediction and remediation workflows.
  • Manage hardware lifecycle and data center operations.
  • Develop services for identity, secrets, tenant configuration, and customer bridging.
Requirements
  • 7+ years of production software engineering experience, including 2 or more years operating what you built (real on‑call experience, not just shipping code).
  • Production‑depth mastery of at least one systems‑grade language—Go (preferred), Rust, or Java.
  • Proficiency in Python for tooling and SDK work.
  • Strong grasp of at‑least‑once vs. exactly‑once trade‑offs, idempotency, back‑pressure, leader election, consistent hashing, gossip, and fan‑out.
  • Experience at production scale with Prometheus, VictoriaMetrics, Mimir, Thanos, Loki, Elasticsearch, Tempo, Jaeger, or OpenTelemetry.
  • Hands‑on experience with Argo, Flux, Helm, Kustomize, Cosign signing, signed‑bundle promotion, and blast‑radius‑aware rollouts.
  • Proven experience writing a controller or CRD handling real production traffic, with a deep understanding of watch‑cache mechanics, leader election, and reconcile loops.
  • Experience executing end‑to‑end mTLS bootstrap with certificate rotation.
  • Hands‑on experience with HashiCorp Vault or cloud KMS (AWS KMS / GCP KMS).
  • Ability to read a Prometheus query plan, build a recording‑rule strategy, and write SQL that joins per‑tenant telemetry against analytics‑lake tables.
  • Rigorous approach to unit, integration, contract, chaos, and soak testing.
  • Ability to author clear design docs that align with existing platform architecture, create runbooks optimized for 3 AM on‑call responses, and write intent‑driven PR descriptions.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior SRE Platform Architect
Senior SRE Platform Architect

Jobtailor • California (MO)

On-site
USD 180,000 - 240,000
SRE / DevOps Engineer
SRE / DevOps Engineer

HeadHR • Town of Poland (NY)

On-site
USD 120,000 - 150,000
Platform SRE Engineer — Build, Run & Scale Systems
Platform SRE Engineer — Build, Run & Scale Systems

Jobtailor • California (MO)

On-site
USD 150,000 - 200,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Madison-Davis, LLC • United States

On-site
USD 120,000 - 150,000
Senior DevOps/SRE Engineer
Senior DevOps/SRE Engineer

HTEC Group • United States

Hybrid
USD 90,000 - 130,000
Associate Engineer, Site Reliability
Associate Engineer, Site Reliability

R&D • United States

On-site
USD 90,000 - 140,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

The ReWork Group • New York (NY)

On-site
USD 120,000 - 160,000
Site Reliability Engineer (SRE)
Site Reliability Engineer (SRE)

Virtual Tech Gurus • Puerto Rico

On-site
USD 140,000 - 210,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Jobtailor • Arlington (VA)

On-site
USD 140,000 - 200,000
Senior Production Engineer
Senior Production Engineer

Anduril Industries • United States

On-site
USD 120,000 - 160,000
Full Family Health Coverage
16 Weeks Paid Parental Leave for All Caregivers
Incentivized Time Off
+1