Staff Site Reliability Engineer

Jobtailor

Town of Florida (NY)

On-site

USD 180,000 - 240,000

Full time

4 days ago
Be an early applicant
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Job summary

Jobtailor is seeking an experienced Platform Architect/Lead to shape major cloud platform decisions and drive reliability across enterprise-scale workloads. You will own design, implementation, and ongoing improvements for critical systems, with heavy emphasis on observability and incident response.

You will lead capacity planning, performance tuning, and security enhancements while guiding on-call readiness and cross-functional collaboration.

Qualifications

  • 5+ years of experience in Cloud Architecture, Site Reliability Engineering, Platform Engineering, or related fields.
  • Proven track record of designing and delivering at enterprise scale.
  • Deep, hands-on expertise with Google Cloud Platform (GCP) and Kubernetes (K8s).
  • Experience running high-volume, high-availability workloads with 99.999% reliability targets.
  • Expert-level command of monitoring, observability, and alerting platforms such as Datadog, Prometheus, Grafana, and PagerDuty.
  • Deep Linux expertise, including kernel internals, system performance tuning, hardening, and production OS troubleshooting.
  • Hands-on experience designing and integrating AI/ML-powered solutions into cloud-native platforms.
  • Familiarity with LLM orchestration, vector databases, model serving infrastructure, and AI observability.
  • Ability to evaluate emerging tools and translate them into reliable, production-grade capabilities.

Responsibilities

  • Contribute to architectural and strategic decisions around major platform components.
  • Lead research, testing, implementation, and continuous improvement for new systems and tooling.
  • Perform capacity planning, load testing, and security improvements.
  • Participate fully in the team’s on-call rotation.
  • Model calm, effective, and blameless incident response.
  • Serve as a significant technical contributor during major incidents and problem resolution.
  • Plan and lead high-risk maintenance events with minimal to no customer impact.
  • Elevate team standards through new tooling, processes, procedures, and effective communication.
  • Lead and represent the team during transitions or coverage gaps.
  • Set professional benchmarks in technical quality, engineering culture, and cross-functional collaboration.
  • Report to a DevOps Director

Skills

Cloud Architecture
Site Reliability Engineering
Google Cloud Platform
Kubernetes
Monitoring & Observability
Linux Expertise
AI/ML Integration
LLM Orchestration
Capacity Planning
Load Testing
Security Improvements
High-Availability Workloads

Tools

Datadog
Prometheus
Grafana
PagerDuty

Job description

  • Contribute to architectural and strategic decisions around major platform components
  • Lead research, testing, implementation, and continuous improvement for new systems and tooling
  • Perform capacity planning, load testing, and security improvements
  • Participate fully in the team’s on-call rotation
  • Model calm, effective, and blameless incident response
  • Serve as a significant technical contributor during major incidents and problem resolution
  • Plan and lead high-risk maintenance events with minimal to no customer impact
  • Elevate team standards through new tooling, processes, procedures, and effective communication
  • Lead and represent the team during transitions or coverage gaps
  • Set professional benchmarks in technical quality, engineering culture, and cross-functional collaboration
  • Report to a DevOps Director
Requirements
  • 5+ years of experience in Cloud Architecture, Site Reliability Engineering, Platform Engineering, or related fields
  • Proven track record of designing and delivering at enterprise scale
  • Deep, hands-on expertise with Google Cloud Platform (GCP) and Kubernetes (K8s)
  • Experience running high-volume, high-availability workloads with 99.999% reliability targets
  • Expert-level command of monitoring, observability, and alerting platforms such as Datadog, Prometheus, Grafana, and PagerDuty
  • Deep Linux expertise, including kernel internals, system performance tuning, hardening, and production OS troubleshooting
  • Hands-on experience designing and integrating AI/ML-powered solutions into cloud-native platforms
  • Familiarity with LLM orchestration, vector databases, model serving infrastructure, and AI observability
  • Ability to evaluate emerging tools and translate them into reliable, production-grade capabilities
Core Competencies

Demonstrates extensive expertise in Cloud Architecture and Site Reliability Engineering, with a strong focus on Google Cloud Platform and Kubernetes. Proven ability to design and implement high-availability systems while leading incident response and continuous improvement initiatives.

Highest-signal resume keywords
  • Cloud Architecture
  • Site Reliability Engineering
  • Google Cloud Platform (GCP)
  • Kubernetes (K8s)
  • Monitoring and Observability
Hard Skills
  • Cloud Architecture
  • Site Reliability Engineering
  • Kubernetes (K8s)
  • Linux Expertise
  • AI/ML Integration
  • Capacity Planning
  • Load Testing
  • Security Improvements
  • High-Availability Workloads
  • Production OS Troubleshooting
Soft Skills
  • Effective Communication
  • Incident Response
  • Cross-Functional Collaboration
  • Leadership
Industry Keywords
  • Enterprise Scale
  • High-Volume Workloads
  • 99.999% Reliability
  • AI Observability
  • LLM Orchestration
Tools & Technologies
  • Datadog
  • Prometheus
  • Grafana
  • PagerDuty
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Site Reliability Engineer NEX
Senior Site Reliability Engineer NEX

NexTier Completion Solutions Inc. • Houston (TX)

On-site
USD 110,000 - 150,000
Senior Site Reliability Engineer – Digital Assets
Senior Site Reliability Engineer – Digital Assets

Jobtailor • Arizona

On-site
USD 120,000 - 170,000
Principal Site Reliability Engineer
Principal Site Reliability Engineer

Jobtailor • Arizona

On-site
USD 180,000 - 240,000
Observability Engineer / Site Reliability Engineer
Observability Engineer / Site Reliability Engineer

Jobtailor • California (MO)

On-site
USD 140,000 - 190,000
Site Reliability Engineer – Lead
Site Reliability Engineer – Lead

Jobtailor • Arizona

On-site
USD 140,000 - 230,000
Staff Site Reliability Engineer, SRE
Staff Site Reliability Engineer, SRE

Jobtailor • California (MO)

On-site
USD 120,000 - 210,000
Systems Engineer - Cloud Ops
Systems Engineer - Cloud Ops

AutoZone • Memphis (TN)

On-site
USD 110,000 - 150,000
Senior Reliability Engineer
Senior Reliability Engineer

Jobtailor • Menomonee Falls (WI)

On-site
USD 120,000 - 160,000
Site Reliability Engineer
Site Reliability Engineer

Compunnel, Inc. • New Jersey

On-site
USD 120,000 - 150,000
Senior Site Reliability Engineer – Google Distributed Cloud Edge (Edge SRE)
Senior Site Reliability Engineer – Google Distributed Cloud Edge (Edge SRE)

CoSourcing Partners Inc. • Chicago (IL)

Hybrid
USD 150,000 - 190,000