Senior Site Reliability Engineer: Scalable, Multi-Cloud

Movable Ink

United States

On-site

USD 184,200 - 240,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Movable Ink seeks a Lead Site Reliability Engineer to own the design and evolution of major systems within a multi‑cloud, active‑active content serving platform that handles billions of requests daily. You will mentor teams, define reliability standards, and drive architectural decisions across infrastructure and software development.

Responsibilites include automation strategy, logging architecture, capacity planning, and cross‑functional reliability initiatives.

Qualifications

  • Proven track record in Site Reliability or Software Engineering, designing, building, and owning scalable, resilient services with a focus on long‑term reliability strategy.
  • Deep expertise in architecting and operating complex distributed systems such as Apache Pulsar, Apache Kafka, Grafana Loki, ScyllaDB/Cassandra, with the ability to guide teams through distributed system challenges.
  • Designing and owning automation strategies to manage services at scale, with expertise in establishing performance analysis frameworks and mentoring others on diagnostics and resolution.
  • Deep, hands‑on experience (6+ years) in Site Reliability or Software Engineering, specifically leading and shaping multi‑cloud architecture and strategy (AWS and GCP).
  • Experience architecting and leading large‑scale observability platforms, including defining observability standards and SLO frameworks. We use Prometheus and Thanos with Grafana Alloy, Loki and Tempo
  • Experience leading on‑call excellence, including driving improvements to monitoring and alerting strategies, automating runbooks and mentoring team members on incident response best practices. Every member of the SRE team does a week long on‑call rotation
  • Expert‑level proficiency with infrastructure as code, including defining IaC standards and patterns across teams. We use Terraform and Chef
  • Advanced Kubernetes expertise, including cluster architecture design, multi‑tenancy strategies, and guiding teams on container orchestration best practices. We use EKS and GKE
  • Proficiency in multiple programming languages with the ability to design and review code that meets reliability standards. We use NodeJS, Golang, Ruby, Python and shell scripting
  • Advanced Linux systems expertise, with the ability to diagnose complex system‑level issues and mentor others on performance tuning and troubleshooting

Responsibilities

  • Define and drive the automation strategy for infrastructure tooling, establishing standards that minimize manual work, increase performance and reduce incident frequency and severity of incidents
  • Own the design, reliability and evolution of core platform applications, mentoring team members on best practices and ensuring systems meet long‑term business objectives
  • Architect and lead the logging platform strategy, driving its design and balancing availability, retention and cost optimization
  • Establish capacity planning and performance management frameworks, proactively identifying scaling opportunities and guiding teams through complex troubleshooting scenarios
  • Lead cross‑functional reliability initiatives with SRE and service engineering teams, influencing architectural decisions and championing practices that ensure resilient service delivery
  • Demonstrate a high level of autonomy in anticipating, identifying, and addressing systemic weaknesses and opportunities for platform improvement without direct supervision.

Skills

Site Reliability
Distributed systems
Observability platforms
IaC automation
Kubernetes
AWS
GCP
NodeJS
Golang
Python
Shell scripting
Linux
Terraform
Chef
On-call

Tools

Grafana Loki
Tempo
Prometheus
Thanos

Job description

Movable Ink seeks a Lead Site Reliability Engineer to own the design and evolution of major systems within a multi‑cloud, active‑active content serving platform that handles billions of requests daily. You will mentor teams, define reliability standards, and drive architectural decisions across infrastructure and software development.

Responsibilites include automation strategy, logging architecture, capacity planning, and cross‑functional reliability initiatives.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Remote Senior Site Reliability Engineer - Cloud & Automation
Remote Senior Site Reliability Engineer - Cloud & Automation

Multi Media LLC • United States

On-site
USD 169,000 - 215,000
Fully Remote
Health Insurance
Vision Insurance
+10
Lead Site Reliability Engineer - Scale & Reliability
Lead Site Reliability Engineer - Scale & Reliability

Doist • United States

Remote
USD 170,000 - 230,000
Lead Site Reliability Engineer
Lead Site Reliability Engineer

Movable Ink • New York (NY)

On-site
USD 184,000 - 240,000
Lead Site Reliability Engineer - Architect & Own Production
Lead Site Reliability Engineer - Architect & Own Production

Optimal Market Technologies • New York (NY)

On-site
USD 175,000 - 200,000
Lead Site Reliability Engineer
Lead Site Reliability Engineer

Movable Ink • United States

On-site
USD 184,000 - 240,000
Senior Site Reliability Engineer – Remote, Impact & Automation
Senior Site Reliability Engineer – Remote, Impact & Automation

Midwest Startups • United States

On-site
USD 175,000 - 185,000
Market-leading medical, dental, and視on
Stock options
Premium-Tier Origin Financial Wellness
+6
Lead Site Reliability Engineer: AWS Cloud & Automation
Lead Site Reliability Engineer: AWS Cloud & Automation

Selby Jennings • Wilmington (NC)

On-site
USD 140,000 - 200,000
Senior Site Reliability Engineer: Scalable Hybrid Infra
Senior Site Reliability Engineer: Scalable Hybrid Infra

Redwood Materials • Nevada (IA)

On-site
USD 140,000 - 180,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Axiom Pursuits • San Francisco (CA)

On-site
USD 150,000 - 180,000
Remote-Optional Senior Site Reliability Engineer
Remote-Optional Senior Site Reliability Engineer

Multi Media, LLC • United States

On-site
USD 169,000 - 215,000
Fully Remote Optional
Health, Vision, Dental, Life Insurance
Unlimited PTO
+4