SRE: 99.9% Uptime in Multi-Cloud & Kubernetes

Gridiron IT

Arlington (VA)

On-site

USD 120,000 - 210,000

Full time

9 days ago
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Benefits offered by this job

Medical insurance
Dental insurance
Vision insurance
HSAs/FSAs
401(k)
Disability & ADD insurance
Life and pet insurance

Job summary

Gridiron IT is seeking a DevOps Site Reliability Engineer (SRE) to ensure continuous availability, performance, and health of the enterprise platform across multi-cloud environments.

You will implement monitoring, alerting, and automation, handle on-call incidents within 1 hour, and collaborate with CSPs to maintain 99.9% uptime. This role requires strong SRE skills, scripting in Python/Bash, and a calm, methodical approach to outages.

Qualifications

  • Strong background in Site Reliability Engineering principles and practices.
  • Hands-on experience with Kubernetes deployments and orchestration.
  • Proven experience managing and troubleshooting multi-cloud environments (GCP, Azure, AWS).
  • Expertise in setting up and managing monitoring, logging, and automated alerting systems.
  • Proficiency in scripting and automation (e.g., Python, Bash) for system stabilization and diagnostic tasks.

Responsibilities

  • Continuous Monitoring: Provide 24/7 infrastructure health monitoring and automated telemetry tracking to maintain the mandated 99.9% core platform availability.
  • Incident Response: Serve on-call to respond to major incidents or platform downtime within 1 hour of notification, executing rapid platform downtime response and system stabilization maneuvers.
  • Automation: Develop and maintain a library of scripts and internal tools to streamline and automate repetitive diagnostic tasks, health checks, and data-gathering procedures.
  • Dashboarding & Telemetry: Create and maintain automated dashboards for uptime, incident status, API latency, and other critical support metrics to ensure platform health visibility.
  • Vendor Coordination: Lead direct engineering-level coordination with Cloud Service Providers (CSPs) during outages or underlying infrastructure issues.
  • Reliability Engineering: Continuously evaluate, monitor, and provide recommended improvements to logging, system metrics, and architecture to ensure the platform remains rapidly scalable and highly available.

Skills

Site Reliability Engineering
Kubernetes
Multi-cloud (GCP/Azure/AWS)
Monitoring & alerting
Scripting (Python/Bash)

Tools

Kubernetes

Job description

Gridiron IT is seeking a DevOps Site Reliability Engineer (SRE) to ensure continuous availability, performance, and health of the enterprise platform across multi-cloud environments.

You will implement monitoring, alerting, and automation, handle on-call incidents within 1 hour, and collaborate with CSPs to maintain 99.9% uptime. This role requires strong SRE skills, scripting in Python/Bash, and a calm, methodical approach to outages.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

SRE DevOps Engineer - 99.9% Uptime, Multi-Cloud
SRE DevOps Engineer - 99.9% Uptime, Multi-Cloud

IT Veterans • Washington

On-site
USD 120,000 - 180,000
Site Reliability Engineer
Site Reliability Engineer

Gridiron IT • Arlington (VA)

On-site
USD 120,000 - 210,000
Medical insurance
Dental insurance
Vision insurance
+4
Senior Site Reliability Engineer: Cloud, Kubernetes Uptime
Senior Site Reliability Engineer: Cloud, Kubernetes Uptime

Compunnel, Inc. • Greenwood Village (CO)

On-site
USD 120,000 - 150,000
Site Reliability Engineer — 99.9% Uptime, Kubernetes
Site Reliability Engineer — 99.9% Uptime, Kubernetes

Latent • San Francisco (CA)

On-site
USD 140,000 - 200,000
Cloud SRE – 24x7 Production Uptime (Hybrid)
Cloud SRE – 24x7 Production Uptime (Hybrid)

Skyhigh Security • Frisco (TX)

Hybrid
USD 110,000 - 140,000
Site Reliability Engineer (SRE)
Site Reliability Engineer (SRE)

Virtual Tech Gurus • Puerto Rico

On-site
USD 140,000 - 210,000
Site Reliability Engineer (SRE)
Site Reliability Engineer (SRE)

myBridge Corporation • Austin (TX)

On-site
USD 120,000 - 160,000
DevOps & SRE — Remote, Multi-Cloud
DevOps & SRE — Remote, Multi-Cloud

AgileEngine, LLC. • United States

On-site
USD 140,000 - 190,000
Growth without limits
Competitive compensation
Flexibility
+3
SRE Leader: Real-Time Healthcare Reliability
SRE Leader: Real-Time Healthcare Reliability

Kontakt Micro-Location Sp. Z.o.o. • New York (NY)

Hybrid
USD 180,000 - 260,000
Equity in a high-growth company
Health, dental, and vision coverage
401k
+3
SRE/DevOps Engineer - 9–5, No On-Call, AWS & Kubernetes
SRE/DevOps Engineer - 9–5, No On-Call, AWS & Kubernetes

INSPYR Solutions • Sunnyvale (CA)

Hybrid
USD 120,000 - 180,000
Work-life balance
No on-call requirements
Standard business hours