Site Reliability Engineer

JobCubby

Barrington, Northern (RI, KY)

Hybrid

USD 110,000 - 170,000

Full time

2 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

JobCubby is seeking an experienced Site Reliability Engineer to design, implement, and maintain scalable, reliable production systems in our cloud-native environment. This role emphasizes automation, observability, and collaboration with software and operations teams to drive uptime and performance improvements.

The ideal candidate will bring strong software engineering skills and hands-on experience with container orchestration, cloud platforms, and modern CI/CD practices.

Qualifications

  • Proven experience as a Site Reliability Engineer, Production Engineer, DevOps Engineer, or similar role.

Responsibilities

  • Design, implement, and maintain highly available production systems.
  • Develop automation and reliability tooling using Go, Python, Java, or Rust.
  • Deploy and manage containerized applications using Docker and Kubernetes.
  • Implement and maintain monitoring, logging, metrics, and distributed tracing solutions.
  • Participate in incident response, root-cause analysis, and post-incident reviews.
  • Collaborate with software developers to improve application reliability and performance.
  • Automate repetitive operational tasks and improve engineering efficiency.
  • Implement Infrastructure as Code using Terraform and/or Ansible.
  • Develop and maintain CI/CD pipelines for reliable deployments.

Skills

Go
Python
Java
Rust
AWS
Azure
GCP
Kubernetes
Docker
Linux
OTel

Tools

Terraform
Ansible
CI/CD
OpenTelemetry
Prometheus
Grafana

Job description

We are seeking an experienced Site Reliability Engineer (SRE) to design, implement, and maintain highly available, scalable, secure, and reliable production systems. The ideal candidate will have strong software engineering skills combined with hands-on experience in cloud infrastructure, Kubernetes, automation, monitoring, and observability.

The SRE will work closely with development, infrastructure, and operations teams to improve system reliability, automate operational processes, and resolve complex production issues.

Requirements
  • Design, deploy, and maintain highly reliable and scalable production systems.
  • Develop automation and reliability tooling using Go, Python, Java, or Rust.
  • Manage and support cloud environments across AWS, Azure, or GCP.
  • Deploy and manage containerized applications using Docker and Kubernetes.
  • Implement and maintain monitoring, logging, metrics, and distributed tracing solutions.
  • Build and enhance observability using OpenTelemetry (OTel) and related technologies.
  • Troubleshoot complex infrastructure, application, and production issues.
  • Participate in incident response, root-cause analysis, and post-incident reviews.
  • Automate repetitive operational tasks and improve engineering efficiency.
  • Implement Infrastructure as Code using Terraform and/or Ansible.
  • Develop and maintain CI/CD pipelines for reliable and automated deployments.
  • Monitor system performance, availability, capacity, and overall reliability.
  • Identify reliability risks and implement proactive solutions.
  • Collaborate with software developers to improve application reliability and performance.
  • Establish and improve SRE best practices, operational procedures, and reliability standards.
Required Skills & Experience
  • Proven experience as a Site Reliability Engineer, Production Engineer, DevOps Engineer, or similar role.
  • Strong programming experience with at least one of:
    • Go/Golang
    • Python
    • Java
    • Rust
  • Hands-hand experience with AWS, Azure, or GCP.
  • Strong experience with Kubernetes and Docker.
  • Strong Linux/Unix administration and troubleshooting skills.
  • Experience with OpenTelemetry and obse-vi-blity.
  • Knowledge of monitoring and visualization tools.
  • Strong knowledge of automation, scripting, networking, and distributed systems.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Site Reliability Engineer (SRE)
Site Reliability Engineer (SRE)

myBridge Corporation • Austin (TX)

On-site
USD 120,000 - 160,000
Site Reliability Engineer
Site Reliability Engineer

TalentDome Staffing • United States

On-site
USD 140,000 - 210,000
Site Reliability Engineer
Site Reliability Engineer

Shya Workforce Solutions • Town of Florida (NY)

On-site
USD 100,000 - 140,000
Site Reliability Engineer
Site Reliability Engineer

SRE • Puerto Rico

Hybrid
USD 120,000 - 180,000
Kubernetes Engineer
Kubernetes Engineer

GCS Recruitment • Maple Shade Township (NJ)

On-site
USD 120,000 - 150,000
Site Reliability Engineer (SRE)
Site Reliability Engineer (SRE)

Virtual Tech Gurus • Puerto Rico

On-site
USD 140,000 - 210,000
Site Reliability Engineer
Site Reliability Engineer

NextGen | GTA: A Kelly Telecom Company • Mount Laurel Township (NJ)

On-site
USD 110,000 - 170,000
Site Reliability Engineer (SRE)
Site Reliability Engineer (SRE)

Methodic • San Francisco (CA)

On-site
USD 140,000 - 210,000
Site Reliability Engineer – Lead
Site Reliability Engineer – Lead

Jobtailor • Arizona

On-site
USD 140,000 - 230,000
Site Reliability Engineer
Site Reliability Engineer

Harrison Clarke • New York (NY)

On-site
USD 120,000 - 160,000