Senior Site Reliability Engineer

Jobtailor

North Carolina

Hybrid

USD 130,000 - 190,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Jobtailor is seeking an experienced Site Reliability Engineer to design, build, and manage our scalable infrastructure across cloud and data center environments. You will drive automation, reliability, and observability for our platform, with a strong focus on Kubernetes and OpenShift.

The role requires hands-on expertise in Python/Go, GitOps workflows, and CI/CD tooling, along with incident response and on-call duties in a fast-moving, collaborative team.

Qualifications

  • 5+ years operating production services on Kubernetes or OpenShift
  • 3+ years programming in Python and Go
  • 2+ years cloud experience (GCP, Azure, AWS)
  • Hands-on Kubernetes/OpenShift, Linux, and AWS
  • Experience with GitOps workflows
  • Strong Linux system administration knowledge (RHEL/Fedora preferred)
  • Understanding of networking and authentication protocols (LDAP)
  • Comfort with incident response and on-call duties
  • Ability to work independently with minimal supervision
  • Knowledge of SRE concepts (SLOs, error budgets, toil)

Responsibilities

  • Design, build, and manage large-scale infrastructure across clouds and data centers
  • Automate cloud infrastructure with autoscaling, load balancing, Python/Go scripting, and monitoring tools
  • Develop OpenShift-related offerings and CI/CD components using OpenShift Pipelines, Tekton, and GitLab
  • Apply GitOps with ArgoCD for declarative platform management
  • Collaborate to break down complex engineering tasks into consumable chunks
  • Design software such as Kubernetes operators, webhooks, and CLI tools
  • Implement monitoring/observability with DataDog, Splunk, Prometheus, Grafana
  • Lead on-call and incident response for high-severity events
  • Mentor peers and share knowledge in an agile team

Skills

Kubernetes/OpenShift
Python
Go
Cloud platforms (GCP/AWS/Azure)
GitOps
Linux administration
Networking
Incident response
SRE principles

Tools

OpenShift
Tekton
GitLab
Prometheus
Grafana
ArgoCD

Job description

Responsibilities
  • Design, build, and manage our large‑scale infrastructure and platform services, including public cloud, private cloud, and datacenter‑based
  • Automate cloud infrastructure through use of technologies such as autoscaling and load balancing, scripting in Python and Golang, and monitoring and alerting solutions such as Splunk, Splunk IM, Prometheus, Grafana, Catchpoint, and DataDog
  • Design, develop, and become an expert in IT’s Red Hat OpenShift offerings by leveraging emerging industry standards
  • Build and support standardized CI/CD platform components using OpenShift Pipelines, Tekton, and GitLab to enable multiple application deployments
  • Apply Infrastructure as Code methodologies using GitOps practices with ArgoCD for declarative platform management
  • Break down complex engineering efforts into consumable chunks while collaborating with teams to understand deliverables
  • Design and develop software such as Kubernetes operators, webhooks, and CLI tools
  • Implement and maintain intelligent infrastructure and application monitoring designed to enable application engineering teams
  • Ensure the production environment operates in accordance with established procedures and best practices
  • Lead escalation support for high‑severity and critical platform‑impacting events
  • Provide feedback on bugs and feature improvements to the various Red Hat Product Engineering teams
  • Design software tests and lead peer reviews to increase the quality of our codebase
  • Help and develop peers’ capabilities through knowledge sharing, mentoring, and collaboration
  • Participate in a regular on‑call schedule, supporting the operation needs of our tenants
  • Drive sustainable incident response and lead blameless postmortems
  • Work within a small agile team to develop and improve SRE methodologies, support peers, plan, and self‑improve
Qualifications
  • 5+ years of experience operating production services on Kubernetes or OpenShift
  • 3+ years of programming experience in Python and Go
  • 2+ years of experience using cloud providers and technologies such as Google Cloud, Azure, and Amazon Web Services
  • Hands‑on experience with Kubernetes/OpenShift, Linux, and AWS
  • Experience with GitOps workflows for managing infrastructure or application configuration
  • Solid understanding of Linux systems administration (RHEL/Fedora preferred)
  • Understanding of standard networking (TCP/IP, DNS, HTTP/TLS) and authentication protocols such as LDAP
  • Comfort with incident response and on‑call responsibilities
  • Ability to work independently with minimal supervision while keeping the team informed
  • Knowledge of SRE principles, including SLOs, error budgets, and toil measurement

Demonstrates expertise in designing and managing large‑scale infrastructure and platform services, with a strong focus on Kubernetes and OpenShift. Proficient in automation, CI/CD practices, and incident response, while fostering collaboration and knowledge sharing within agile teams.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Site Reliability Engineer Lead (SRE) - Internal Kubernetes Container Platform (IKCP)
Site Reliability Engineer Lead (SRE) - Internal Kubernetes Container Platform (IKCP)

Koitecc Solutions • Chandler (AZ), Northern (KY)

Hybrid
USD 140,000 - 200,000
Site Reliability Engineer Lead (SRE) – Internal Kubernetes Container Platform (IKCP)
Site Reliability Engineer Lead (SRE) – Internal Kubernetes Container Platform (IKCP)

Hobbsnews • Chandler (AZ), Northern (KY)

Hybrid
USD 120,000 - 180,000
Site Reliability Engineer Lead (SRE) – Internal Kubernetes Container Platform (IKCP)
Site Reliability Engineer Lead (SRE) – Internal Kubernetes Container Platform (IKCP)

Bank of America • Chandler (AZ)

On-site
USD 140,000 - 170,000
Principal, Infrastructure & Automation (Linux, Kubernetes, Terraform, Ansible, OpernShift)
Principal, Infrastructure & Automation (Linux, Kubernetes, Terraform, Ansible, OpernShift)

Request Technology, LLC • Chicago (IL)

Hybrid
USD 180,000 - 240,000
Site Reliability Engineer (SRE)
Site Reliability Engineer (SRE)

JPS Tech Solutions • Colorado

On-site
USD 160,000 - 230,000
Site Reliability Engineer – Lead
Site Reliability Engineer – Lead

Jobtailor • Arizona

On-site
USD 140,000 - 230,000
Senior OpenShift SRE: Scale Hybrid Cloud Platforms
Senior OpenShift SRE: Scale Hybrid Cloud Platforms

Red Hat, Inc. • Raleigh (NC)

Hybrid
USD 118,000 - 196,000
Medical, dental, vision
401(k) with match
Paid time off
+2
Site Reliability Engineer
Site Reliability Engineer

SCIGON • Naperville (IL)

Hybrid
USD 110,000 - 170,000
Senior SRE — Cloud Infra, Kubernetes & OpenShift
Senior SRE — Cloud Infra, Kubernetes & OpenShift

Jobtailor • North Carolina

Hybrid
USD 130,000 - 190,000
Senior SRE Engineer
Senior SRE Engineer

Compunnel, Inc. • Alpharetta (GA)

On-site
USD 140,000 - 190,000