Senior Site Reliability Engineer / DevOps Engineer

Prophet Town

Mountain View (CA)

On-site

USD 130,000 - 180,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Prophet Town is seeking a Senior Site Reliability Engineer / DevOps Engineer in Mountain View, California. This role requires 5+ years of experience managing global infrastructure, with a focus on Kubernetes, Terraform, and AWS. Candidates will design and operate resilient production systems, tackle cross-region challenges, and collaborate closely with platform engineering. The position offers significant ownership of foundational infrastructure decisions, working with cutting-edge technologies in a dynamic, hands-on environment.

Qualifications

  • 5+ years in Site Reliability Engineering, DevOps, or Platform Engineering.
  • Deep production experience with ArgoCD, Terraform, AWS, and CI/CD systems.
  • Strong networking knowledge, including BGP, VPNs, and DNS.

Responsibilities

  • Design and operate globally distributed production infrastructure.
  • Build highly available multi-region systems with disaster recovery strategies.
  • Manage GitOps workflows and promote pipelines across environments.

Skills

Kubernetes
ArgoCD
Terraform
CI/CD
AWS
Networking

Job description

Mountain View, United States | Posted on 05/12/2026

Location: Onsite - Mountain View, CA

Experience Required: 5+ years

Infrastructure Footprint: Global production infrastructure across AWS, South America, and Europe

Role Overview

Seeking a Senior Site Reliability Engineer / DevOps Engineer to design, scale, and operate highly available global infrastructure supporting production systems across multiple international regions.

This role is for an engineer with 5+ years of experience building and running production‑grade cloud infrastructure. The right person understands where distributed systems fail and has learned the hard lessons that come from operating Kubernetes and cloud platforms at scale.

The ideal candidate has deep hands‑on experience with Kubernetes, ArgoCD, Terraform, CI/CD pipelines, AWS infrastructure, and multi‑region platform reliability. They should understand the limitations, sharp edges, and operational failure modes of these tools.

This is an onsite role working closely with platform engineering and leadership to build resilient global infrastructure.

What You’ll Do
  • Design and operate globally distributed production infrastructure across AWS regions and physical data center environments in South America and Europe
  • Build highly available multi-region systems with strong disaster recovery and failover strategies
  • Solve cross-region networking, latency, DNS routing, replication, and reliability challenges
  • Build, scale, secure, and troubleshoot production Kubernetes clusters
  • Handle cluster lifecycle management, upgrades, node failures, networking issues, storage problems, and control‑plane troubleshooting
  • Tune workloads for resiliency, scheduling efficiency, autoscaling behavior, and resource optimization
  • etcd instability
  • networking overlays and CNI failures
  • node pressure and eviction behaviorcluster upgrade regressions
GitOps / ArgoCD Operations
  • Design and maintain GitOps workflows using ArgoCD
  • Manage promotion pipelines across environments and regions
  • Resolve drift detection issues, sync conflicts, reconciliation failures, and deployment ordering challenges
  • Build safe rollback and progressive deployment strategies

Candidates should know why ArgoCD breaks, not just how to click “Sync.”

Infrastructure as Code
  • Build and maintain reusable Terraform modules for multi‑region infrastructure
  • Manage state strategy, workspace isolation, secrets handling, and provider complexity
  • Solve real‑world Terraform pain points, including:
    • state corruption and locking conflicts
    • module version drift
    • provider upgrade regressions
    • dependency graph surprises
    • cross‑account provisioning complexity
  • Build and optimize production CI/CD pipelines
  • Improve deployment speed, safety, and repeatability
  • Troubleshoot flaky pipelines, artifact inconsistencies, race conditions, environment drift, and rollback failures
Reliability & Observability
  • Establish SLIs/SLOs and production health standards
  • Build alerting, monitoring, tracing, and incident response workflows
  • Lead root cause analysis and postmortem improvements
  • Reduce operational toil through automation
Why This Role

You’ll own foundational infrastructure decisions for globally distributed systems and help build resilient platform capabilities at international scale.

This is a hands‑on engineering role for someone who wants meaningful ownership and complex technical problems.

Requirements
Required Experience

5+ years in Site Reliability Engineering, DevOps, or Platform Engineering

Deep production experience with:

ArgoCD

Terraform

AWS

CI/CD systems

Preferred Experience
  • Experience operating infrastructure across multiple continents
  • Experience with hybrid cloud or physical data center integration
  • Strong networking knowledge, including BGP, VPNs, routing, DNS, and load balancing
  • Experience with security hardening and compliance in production systems
  • Software engineering background with Go, Python, or Bash
What “Senior” Means Here

You have enough production experience to have strong opinions because you have seen failures firsthand.

You know:

why Terraform plans sometimes lie

why ArgoCD syncs can fail for non‑obvious reasons

why Kubernetes upgrades can ruin your week

why “works in staging” means very little

why multi‑region failover diagrams often fail in production

why observability usually breaks exactly when needed most

You’ve solved these problems repeatedly and improved systems because of those lessons.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Site Reliability Engineer
Senior Site Reliability Engineer

Veloc Inc • Coppell (TX)

On-site
USD 140,000 - 190,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Circle Internet Financial, LLC • United States

On-site
USD 140,000 - 190,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Jobtailor • Arlington (VA)

On-site
USD 140,000 - 200,000
Site Reliability Engineer
Site Reliability Engineer

Harrison Clarke • New York (NY)

On-site
USD 120,000 - 160,000
Site Reliability Engineer
Site Reliability Engineer

TalentDome Staffing • United States

On-site
USD 140,000 - 210,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

MeridianLink, Inc. • Northern (KY)

Hybrid
USD 120,000 - 170,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

PayNearMe MT, Inc • United States

On-site
USD 130,000 - 170,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

The ReWork Group • New York (NY)

On-site
USD 120,000 - 160,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

DevOpsChat • United States

Remote
USD 100,000 - 130,000
Competitive salary
Flexible working hours
Professional development opportunities
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Mission Staffing • New York (NY)

Hybrid
USD 140,000 - 200,000