Senior Site Reliability Engineer

ARA

Albuquerque (NM)

Hybrid

USD 120,000 - 180,000

Full time

14 hours ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

ARA is seeking a seasoned Site Reliability Engineer to partner with software developers, platform engineers, and IT staff to enhance system design, operability, and deployment safety. The role focuses on ensuring production support readiness, reliability, and ongoing platform stability in enterprise environments.

Candidates should have 8+ years in SRE/DevOps or related roles, strong Linux and on-prem Kubernetes experience, and proficiency with Helm, Kustomize, FluxCD/ArgoCD, and

Qualifications

  • 8+ years in SRE/DevOps, Platform Eng, or related infrastructure roles supporting production services.
  • Strong Linux administration and troubleshooting in enterprise environments.
  • Experience operating on-prem Kubernetes platforms with CRI, CNI, CSI plugins.
  • Experience deploying apps on Kubernetes using Helm and Kustomize.
  • Experience with GitLab, Artifactory, Jira, Confluence.
  • Experience with GitOps tools FluxCD or ArgoCD.
  • Proficiency scripting with Python, Go, or Bash.
  • Strong observability tooling: monitoring, dashboards, logging, tracing, and SLOs.
  • Understanding reliability engineering concepts: service health indicators, HA, incident response.
  • Must be able to obtain a security clearance (U.S. citizenship).

Responsibilities

  • Partner with software developers, platform engineers, and IT staff to improve system design, operability, deployment safety, and production support readiness.
  • Define and maintain operational standards, runbooks, support procedures, escalation paths, and service-level objectives.
  • Evaluate system architecture to balance functional requirements, service quality, reliability, security, and compliance needs.
  • Drive continuous improvement in platform stability, maintenance, and availability.
  • Provide advanced technical support and troubleshooting for complex platform and service issues affecting internal users and stakeholders.

Skills

Linux administration
Kubernetes
Helm & Kustomize
GitOps FluxCD ArgoCD
Scripting Python Go Bash
Observability tooling
Reliability engineering concepts
Security clearance eligibility

Education

Bachelor's degree in CS or related IT field

Tools

GitLab
Artifactory
Jira
Confluence
FluxCD
ArgoCD

Job description

Essential Functions
  • Partner with software developers, platform engineers, and IT staff to improve system design, operability, deployment safety, and production support readiness.
  • Define and maintain operational standards, runbooks, support procedures, escalation paths, and service-level objectives.
  • Evaluate system architecture and changes to ensure they balance functional requirements, service quality, reliability, security, and compliance needs.
  • Drive continuous improvement in platform stability, maintenance, and availability.
  • Provide advanced technical support and troubleshooting for complex platform and service issues affecting internal users and stakeholders.
  • Partner with software developers, platform engineers, and IT staff to improve system design, operability, deployment safety, and production support readiness.
  • Define and maintain operational standards, runbooks, support procedures, escalation paths, and service-level objectives.
  • Evaluate system architecture and changes to ensure they balance functional requirements, service quality, reliability, security, and compliance needs.
  • Drive continuous improvement in platform stability, maintenance, and availability.
  • Provide advanced technical support and troubleshooting for complex platform and service issues affecting internal users and stakeholders.
Experience And Skills Required
  • 8+ years of experience in Site Reliability Engineering, DevOps, Platform Engineering, Systems Engineering, or related infrastructure roles supporting production services.
  • Strong experience with Linux systems administration and troubleshooting in enterprise environments.
  • Strong experience operating and maintaining on-prem Kubernetes platforms and all related components including CRI, CNI, and CSI plugins.
  • Experience deploying and maintaining applications on Kubernetes using Helm, Kustomize, and similar tooling.
  • Experience supporting DevOps tooling such as GitLab, Artifactory, Jira, Confluence.
  • Experience with GitOps tools such as FluxCD or ArgoCD.
  • Proficiency scripting with at least one of Python, Go, or Bash.
  • Strong experience designing, maintaining, and maturing observability tooling including monitoring, dashboards, logging and tracing, and supporting SLOs.
  • Strong understanding of reliability engineering concepts:
    • Service health indicators
    • High availability design, failure reduction, and testing
    • Operational readiness practices, including developing documentation, runbooks, and architectural descriptions
    • Incident response, root cause analysis, remediation/recovery
  • Ability to obtain a security clearance, which includes U.S. citizenship.
Preferred
  • Experience with multiple Linux distributions including Ubuntu.
  • Experience with at least one of the following: Tanzu Kubernetes, Nutanix Kubernetes Platform, Canonical Kubernetes.
  • Experience with cloud platforms such as AWS and Azure.
  • Experience with infrastructure automation and configuration management.
  • Experience managing AI tooling on Kubernetes including MCP Servers, LLM platforms (vLLM, Ollama), Kubeflow.
  • Experience with security and compliance considerations in regulated environments.
  • DoD experience.
  • Active or inactive Secret Security Clearance.
Education
  • Bachelor's degree in CS, Software Engineering or other IT-related field or equivalent experience
REMOTE WORK NOTICE:

This position may be performed fully remote, hybrid, or onsite at an ARA office. Preference will be given to candidates located onsite in the Albuquerque area.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Site Reliability Engineer
Senior Site Reliability Engineer

Jobtailor • Arlington (VA)

On-site
USD 140,000 - 200,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Veloc Inc • Coppell (TX)

On-site
USD 140,000 - 190,000
Site Reliability Engineer
Site Reliability Engineer

TalentDome Staffing • United States

On-site
USD 140,000 - 210,000
Site Reliability Engineer Lead (SRE) – Internal Kubernetes Container Platform (IKCP)
Site Reliability Engineer Lead (SRE) – Internal Kubernetes Container Platform (IKCP)

Hobbsnews • Chandler (AZ), Northern (KY)

Hybrid
USD 120,000 - 180,000
Site Reliability Engineer Lead (SRE) – Internal Kubernetes Container Platform (IKCP)
Site Reliability Engineer Lead (SRE) – Internal Kubernetes Container Platform (IKCP)

Bank of America • Chandler (AZ)

On-site
USD 140,000 - 190,000
Site Reliability Engineer
Site Reliability Engineer

Jobtailor • California (MO)

Hybrid
USD 120,000 - 160,000
Platform Site Reliability Engineer
Platform Site Reliability Engineer

Specter • San Francisco (CA), Northern (KY)

Hybrid
USD 180,000 - 250,000
Site Reliability Engineer
Site Reliability Engineer

Request Technology, LLC • Chicago (IL)

Hybrid
USD 150,000 - 155,000
Site Reliability Engineer
Site Reliability Engineer

Govcio LLC • Arlington (TX)

Hybrid
USD 230,000 - 250,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

O.C. Tanner • Salt Lake City (UT)

On-site
USD 130,000 - 180,000