Senior Site Reliability Engineer

platform.sh

United States

Remote

USD 120,000 - 190,000

Full time

3 days ago
Be an early applicant
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Job summary

Upsun (formerly Platform.sh) is seeking a Senior Site Reliability Engineer to lead the evolution of our cloud application platform from traditional operations to automation-driven SRE.

You will own engineering workstreams that improve reliability, scalability, and efficiency across multi-cloud environments, partnering with engineering, product, and platform teams to bake reliability into the software delivery lifecycle.

Qualifications

  • 5+ years of SRE or platform engineering experience.
  • Experience with multi-cloud environments (AWS, GCP, Azure) and IaC.
  • Strong observability and incident response skills.

Responsibilities

  • Drive reliability observability strategy with SLIs/SLOs.
  • Automate infrastructure workflows with Terraform and Ansible across clouds.
  • Scale CI/CD delivery pipelines for fast, secure releases.
  • Lead incident response post-mortems and drive blameless analysis.

Skills

SRE practices
Prometheus
Grafana
ELK Stack
Terraform
Ansible
AWS
GCP
CI/CD

Tools

Terraform
Ansible
Prometheus
Grafana
ELK

Job description

About Upsun (formerly Platform.sh)

Upsun is the software factory for AI-human workflows. It is built for today’s hybrid teams, where AI agents write and test code and humans focus on solving the problems that really matter. Developers, DevOps engineers, and platform teams use Upsun to build, ship, and scale confidently without wrestling with backend infrastructure. We give you your time back. You get:

  • Predictable performance, even at scale
  • Secure, compliant environments by default
  • Real-time observability and profiling built in
  • Cloning, configuration, and provisioning in seconds
  • AI-ready features that plug directly into your stack

The name says it all. "Up" means uptime, reliability, and acceleration. "Sun" reflects our follow-the-sun-support, a 24x7, globally distributed support team keeping the lights on while you rest. Our core belief is that software should power brighter solutions and greater innovation.

Upsunners are a remote, global workforce, and we thrive in a multicultural team. We are committed to open source and an open, welcoming environment. Our team spans the globe and the experience spectrum.

What's our commonality, our cultural fabric? A curious spirit and a thirst for knowledge; an eagerness for innovative ideas and cultures. We believe we can build anything together in an environment that frees you to do your best work.

The values:

  • We make a positive impact.
  • We aim for the stars.
  • We care for each other.
Impact of a Senior Site Reliability Engineer

As a Senior Site Reliability Engineer at Upsun, you will lead the evolution of our cloud application platform from traditional cloud operations into a proactive, automation-driven SRE model. You will own critical engineering workstreams that enhance system reliability, scalability, and operational efficiency across multi-cloud environments. Partnering closely with engineering, product, and platform teams, you will embed reliability and performance into every stage of the software delivery lifecycle. In this role, you will anticipate architectural bottlenecks, drive infrastructure-as-code practices, and establish robust observability standards that ensure long-term system stability and uptime for our global users.

What to expect
  • Drive reliability observability strategy: Architect and elevate system monitoring, alerting, and logging using Prometheus, Grafana, and ELK Stack, establishing actionable SLIs/SLOs aligned with core business metrics.
  • Automate infrastructure workflows: Eliminate operational toil by designing and implementing resilient, automated solutions using IaC tools like Terraform and Ansible across AWS, GCP, and Azure.
  • Scale CI/CD delivery pipelines: Optimize pipeline architectures for fast, secure, and zero-downtime releases, ensuring infrastructure resilience during high-volume deployment cycles.
  • Lead incident response post-mortems: Guide high-priority incident triage, drive blameless post-mortem analysis, and implement preventative me
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior SRE - Cloud Reliability & Automation (Remote)
Senior SRE - Cloud Reliability & Automation (Remote)

platform.sh • United States

Remote
USD 120,000 - 190,000
Senior Software Engineer
Senior Software Engineer

platform.sh • United States

Remote
USD 140,000 - 210,000
Senior Analytics Engineer
Senior Analytics Engineer

Webhosting • Charlotte (NC), Northern (KY)

On-site
USD 120,000 - 180,000
Product Manager
Product Manager

Upsun (via Remote Woman) • California (MO)

On-site
USD 120,000 - 180,000
Comprehensive healthcare coverage (UK,
Company stock options
Site Reliability Engineer Engineer
Site Reliability Engineer Engineer

Modus Create • Aurora (IL)

Remote
USD 120,000 - 160,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Alembic • Dunwoody (GA)

On-site
USD 150,000 - 190,000
Director, Sales
Director, Sales

Webhosting • United States

On-site
USD 180,000 - 250,000
Remote-friendly
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Satsuma AI, Inc. • Austin (TX), Northern (KY)

On-site
USD 140,000 - 210,000
Unlimited PTO
401(K)
Healthcare Stipend
+1
Senior Site Reliability Engineer
Senior Site Reliability Engineer

SDI International • Chicago (IL)

On-site
USD 130,000 - 180,000
Senior SRE: AI-Driven Cloud Reliability
Senior SRE: AI-Driven Cloud Reliability

SupportFinity™ • San Francisco (CA)

Hybrid
USD 164,000 - 205,000
BetterUp coaching
Competitive pay
Medical, dental, and vision insurance
+7