Senior Site Reliability Engineer

Runware

United States

Remote

USD 140,000 - 190,000

Full time

14 days+
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Benefits offered by this job

Generous paid time off
Stock options
Remote-first setup
Flexible hours
Family leave
Company retreats

Job summary

Runware is seeking a Site Reliability Engineer to ensure reliability and performance of production services. You will join a remote-first team and work across software, infrastructure and production operations to reduce toil and improve observability and resilience.

You will own incident response, capacity planning and automation, collaborating with engineering and DevOps as the platform scales. Strong Kubernetes, containers and IaC experience required.

Qualifications

  • Experience operating production systems at scale (SRE/Production Engineering/Platform Engineering).
  • Strong understanding of distributed systems across apps, databases, queues, containers, networking and infra.
  • Experience with Kubernetes, containers, IaC and automated deployment practices, plus scripting in Python/Go/PHP.

Responsibilities

  • Own and improve reliability, availability and performance of critical production services across the Runware platform.
  • Define and evolve reliability practices: SLIs/SLOs, alerting, observability and production-readiness.
  • Investigate complex production issues across distributed systems and participate in on-call rotation.
  • Lead incident reviews and RCAs to drive lasting engineering improvements.
  • Reduce toil via automation and safer deployment/recovery processes.
  • Collaborate with Engineering/DevOps on capacity planning, scaling and architecture.

Skills

Kubernetes
Containers
Python
Go
PHP
IaC
Distributed systems
On-call rotation
Observability
MySQL
Redis
RabbitMQ
ClickHouse
Go/or Python/ PHP

Tools

Docker
CI/CD

Job description

Runware is building high-performance infrastructure and products to power the worlds intelligence. Our platform enables developers and businesses to run fast, scalable inference across image, video and emerging modalities, while our Serverless platform allows customers to deploy and scale their own AI models on production-grade GPU infrastructure.

As a Site Reliability Engineer at Runware, you will help ensure these systems remain reliable, performant and resilient as we scale. This is a highly technical, hands-on role working across software, infrastructure and production operations to improve observability, reduce incidents, eliminate operational toil and build lasting improvements across complex distributed systems.

What you’ll do
  • Own and improve the reliability, availability and performance of critical production services across the Runware platform
  • Define and evolve our reliability practices, including SLIs, SLOs, alerting, observability and production-readiness standards
  • Investigate complex production issues across distributed systems, APIs, networking, queues, databases and GPU-backed workloads, participating in our engineering on-call rotation
  • Lead and contribute to incident reviews and RCAs, turning recurring failure modes into lasting engineering improvements
  • Reduce operational toil through automation, automated remediation and improvements to deployment safety, recovery and system resilience
  • Work closely with Engineering and DevOps teams on capacity planning, performance, scaling and architectural improvements as the platform grows
  • Have strong experience operating and troubleshooting production systems at scale in an SRE, Production Engineering, Platform Engineering or similar role
  • Have a strong understanding of distributed systems and are comfortable debugging across applications, databases, queues, containers, networking and infrastructure
  • Have experience designing and operating observability systems using metrics, logs and distributed tracing
  • Understand SRE principles including SLIs, SLOs, error budgets, capacity planning, incident management and reducing operational toil
  • Have experience with Kubernetes, containers, IaC and automated deployment practices, alongside the ability to write software and automation using languages such as Python, Go or PHP
  • Take strong ownership of production problems and are comfortable participating in an engineering on-call rotation, taking issues from initial investigation through to long-term remediation
  • Experience operating high-throughput or low-latency APIs and distributed systems
  • Experience with bare-metal infrastructure, GPU environments or AI and ML workloads
  • Experience with RabbitMQ or other distributed messaging and queueing systems
  • Experience operating MySQL, Redis, ClickHouse or similar production data systems
  • Experience with global traffic management, load balancing, CDN platforms and hybrid infrastructure environments
  • Experience building automated scaling, capacity management or self-healing systems

We're a remote-first collective, meeting in person twice a year to plan, brainstorm, celebrate wins, and enjoy some face-to-face time. We have core hours for cooperative working and calls, but outside of that your calendar is yours. Work the hours that let you perform at your peak while also building a healthy life.

Our release cycles are fast and intense, but they're followed by real downtime. After big pushes we expect the team to unplug, recharge, and come back ready & stronger than ever for the next leap.

  • Generous paid time off - vacation, sick days, public holidays
  • Meaningful stock options - share in the upside you create
  • Remote-first setup - work from home anywhere we can employ you
  • Flexible hours - own your schedule outside core collaboration blocks
  • Family leave - paid maternity, paternity, and caregiver time
  • Company retreats - twice-yearly gatherings in inspiring locations
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Remote Senior SRE: Build Reliable, Scalable AI Infra
Remote Senior SRE: Build Reliable, Scalable AI Infra

Runware • Town of Sweden (NY)

On-site
USD 140,000 - 190,000
Generous paid time off
Meaningful stock options
Remote-first setup
+3
Staff Software Engineer - Serverless
Staff Software Engineer - Serverless

Runware • United States

On-site
USD 180,000 - 240,000
Generous paid time off
Meaningful stock options
Remote-first setup
+3
Senior Software Engineer
Senior Software Engineer

Runware • Northern (KY)

On-site
USD 140,000 - 210,000
Generous stock options
Remote-first setup
Flexible hours
+1
Senior SRE: Remote-First, Flexible Hours, Stock Options
Senior SRE: Remote-First, Flexible Hours, Stock Options

Runware • United States

Remote
USD 140,000 - 190,000
Generous paid time off
Stock options
Remote-first setup
+3
Founding Engineer - Site Reliability
Founding Engineer - Site Reliability

uRun • San Francisco (CA)

On-site
USD 120,000 - 140,000
Health, dental, and vision – full coverage
401(k) – company-supported retirement savings
FSA/HSA – flexible spending accounts
+3
Staff Site Reliability Engineer
Staff Site Reliability Engineer

Wand AI • Palo Alto (CA)

On-site
USD 180,000 - 250,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Justjoin • United States

Remote
USD 140,000 - 190,000
Health benefits
Financial planning
Family benefits
+2
Engineering Manager
Engineering Manager

Runware • United States

On-site
USD 140,000 - 210,000
Generous PTO
Stock options
Remote-first
+3
Senior Site Reliability Engineer
Senior Site Reliability Engineer

The ReWork Group • New York (NY)

On-site
USD 120,000 - 160,000
Developer Relations Engineer
Developer Relations Engineer

Runware • San Francisco (CA)

On-site
USD 100,000 - 140,000
Generous paid time off
Meaningful stock options
Remote-first setup
+3