Senior SRE: AI-First Platform & Observability Lead (Remote)

Bot Jobs

Myrtle Point (OR)

Remote

USD 170,000 - 210,000

Full time

13 days ago
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Benefits offered by this job

In-person connection offsites
Tech & learning stipend
Remote by design
Health & wellness benefits
Equity with upside

Job summary

Replicant is seeking a Senior Site Reliability Engineer to help design and operate an AI-native platform with a focus on reliability, scale, and secure cloud infrastructure. You will own CI/CD, observability, and on-call processes while collaborating across fully remote teams.

The role emphasizes platform engineering, incident management, and developer self-service in a Kubernetes/GCP environment, with a preference for strong experience in Node/TypeScript, Python, and Terraform.

Qualifications

  • 6+ years of experience in software development enablement roles.
  • Solid experience owning CI/CD platforms end to end, including caching, architecture, and developer self-service.
  • Effective use of AI tools for coding, troubleshooting, and reasoning with a defensible stance on AI usage.
  • Familiarity with Node/TypeScript and Python, plus Terraform for automation, in a Kubernetes/Helm ecosystem.
  • Practical experience with observability: logs, metrics, tracing, monitoring, and incident management.

Responsibilities

  • Contribute to patterns, design, and implementation of platform domains to shape Replicant's future engineering.
  • Build and improve systems to reduce toil and keep production infra available under large-scale real-time AI traffic.
  • Extend and iterate the agent harness with CI, sandboxes, guardrails, and evaluation loops.
  • Own and improve CI/CD pipelines and developer tooling for performance, deployment ergonomics, and new services.
  • Participate in on-call rotation and incident management to ensure platform uptime.

Skills

CI/CD pipelines
Node/TypeScript
Python
Terraform
Kubernetes/Helm
Observability
Remote teamwork
AI tools usage

Tools

Datadog
Prometheus
Grafana
GitLab CI
Helm
Terraform
Kubernetes

Job description

Replicant is seeking a Senior Site Reliability Engineer to help design and operate an AI-native platform with a focus on reliability, scale, and secure cloud infrastructure. You will own CI/CD, observability, and on-call processes while collaborating across fully remote teams.

The role emphasizes platform engineering, incident management, and developer self-service in a Kubernetes/GCP environment, with a preference for strong experience in Node/TypeScript, Python, and Terraform.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior SRE for AI-First Platform & Scale
Senior SRE for AI-First Platform & Scale

Replicant • United States

On-site
USD 140,000 - 210,000
Offsites
Tech & learning stipend
Remote by design
+3
Remote SRE: AI Platform Reliability & Automation
Remote SRE: AI Platform Reliability & Automation

Runpod • United States

On-site
USD 150,000 - 200,000
Remote work first
Competitive base salary
Stock options equity
+2
Remote Senior SRE - AI Infra, Kubernetes & Terraform
Remote Senior SRE - AI Infra, Kubernetes & Terraform

Motion Recruitment • United States

Remote
USD 140,000 - 170,000
Medical, dental, and vision
Equity / Stock Options
Remote equipment stipend
+3
Senior Site Reliability Engineer (Fully Remote)
Senior Site Reliability Engineer (Fully Remote)

Replicant • United States

Remote
USD 120,000 - 180,000
Senior SRE for AI-Native Platform — Remote US/Canada
Senior SRE for AI-Native Platform — Remote US/Canada

MAP SSG • United States

Remote
USD 160,000 - 210,000
Equity compensation
Senior SRE: Scale Resilient AI Platforms & Automation
Senior SRE: Scale Resilient AI Platforms & Automation

Relx Plc • Philadelphia

Hybrid
USD 95,000 - 159,000
Senior SRE - Remote Observability for AI Infrastructure
Senior SRE - Remote Observability for AI Infrastructure

Cribl • Atlanta (GA)

Remote
USD 142,000 - 195,000
health
dental
vision
+7
Senior SRE: Observability, Automation & Scalable Systems
Senior SRE: Observability, Automation & Scalable Systems

Replit • Northern (KY)

Hybrid
USD 140,000 - 190,000
Competitive Salary & Equity
401(k) 4% match (US)
Health, Dental, Vision & Life
+7
Remote Senior SRE: Platform Reliability & Incidents
Remote Senior SRE: Platform Reliability & Incidents

Tamarind Intelligence • United States

Remote
USD 99,000 - 140,000
Health coverage for you and dependents
Tech spending stipend
Employee stock purchase plan
Staff SRE: Scale, Observability & Automation Leader
Staff SRE: Scale, Observability & Automation Leader

Replit • Northern (KY)

Hybrid
USD 180,000 - 260,000
Salary & equity
401(k) matching
Health, dental, vision, life
+9