Reliability Engineer

Compunnel, Inc.

Town of Texas (WI)

On-site

USD 140,000 - 190,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Compunnel, Inc. is seeking a Reliability Engineer to help scale production insight, data, and backup recovery, with a focus on automation and observability across container platforms in a cloud environment.

You will drive recovery automation, manage Kubernetes clusters, and collaborate with client teams to improve resiliency and incident response across distributed systems.

Qualifications

  • Bachelor’s Degree or equivalent experience in a technology related field is required.
  • Production experience running Cloud and on‑prem Storage workloads at scale.
  • Experience managing and maintaining Kubernetes Clusters on EKS/AKS and RKS.
  • Drive for continuous improvement and ability to tackle complex problems.
  • Experience managing and interpreting large datasets using query languages and visualization tools (PowerBI/Tableau).
  • Experience in software development with Python, NodeJS, or Java with a focus on SDLC and automation.
  • 5–7 years of hands‑on experience deploying and/or supporting highly distributed multi‑tiered systems at scale.
  • Experience building and deploying Docker images including Docker Compose.
  • Hands‑on experience with Jenkins Core, including authoring and maintaining declarative CI/CD pipelines and libraries.
  • Experience with distributed version control systems, Git preferred.
  • Experience crafting and maintaining logging, monitoring, and alerting using Datadog and Splunk.
  • Practical experience in building cloud hosted and native applications for the enterprise.
  • Deep understanding of AWS/Azure services that support reliability, observability, and automation/orchestration.
  • Experience in incident/crisis management and supporting critically important applications.
  • Experience with observability tools (Prometheus, Grafana, ELK/OpenSearch, OpenTelemetry, Datadog).
  • Ability to automate with scripting languages (Python, Shell scripting, etc.).
  • Experience managing systems using infrastructure as code tools (IAM, ARM, Terraform, Chef).

Responsibilities

  • Automate recovery workflows.
  • Rehouse recovery into alternate datacenters.
  • Test recovery processes.
  • Advance enterprise resiliency through improved recovery capabilities.
  • Reduce recovery time via automation.
  • Enable rehoused recovery into new datacenters.
  • Strengthen platform reliability through data protection design.
  • Craft scalable solutions and automation to monitor the health and establish signals to drive understanding of Container Platform environments.
  • Strengthen operational processes with client’s support and incident management teams for cloud ecosystem.
  • Drive ongoing reliability improvements in their Kubernetes service offerings.
  • Prioritize client / partner needs to serve as their voice and guide execution of the team.

Skills

Kubernetes
Python
NodeJS
Java
Git
Terraform
Jenkins CI/CD
AWS/Azure
OpenTelemetry
Datadog
Splunk
Docker
CI/CD pipelines

Education

Bachelor’s Degree or equivalent experience in a technology related field

Tools

Kubernetes (EKS/AKS/RKS)
Docker
Jenkins Core
Git
Datadog
Splunk
OpenTelemetry
Terraform
IAM/ARM/Terraform/Chef
PowerBI/Tableau

Job description

JOB SUMMARY

The Reliability Engineering group within Enterprise Infrastructure combines Operations Excellence with the Development Experience to deliver services at high scale, high availability with resilience by using automation and Infrastructure Code. We build reliability into our ecosystem by applying standards in Resiliency Engineering, Automation, Observability & Chaos Testing. Additionally, this role contributes to enterprise backup and recovery capabilities including automation of recovery workflows, rehoused recovery into alternate datacenters, and testing of recovery processes. We are looking for a systems thinking, reliability engineer who has helped teams scale through production insight, data and backup recovery, operational automation, developer guidance, real-time metrics, automation, automation, automation. Crafting scalable solutions and automation to monitor the health and establish signals to drive understanding of our Container Platform environments. Strengthening operational processes with client’s support and incident management teams for our cloud ecosystem Working with client’s and cloud service provider product teams and driving ongoing reliability improvements in their Kubernetes service offerings. Anticipating, discovering through ongoing interaction with, and prioritizing client / partner needs to serve as their voice and guide execution of the team.

Key Responsibilities
  • Automate recovery workflows
  • Rehouse recovery into alternate datacenters
  • Test recovery processes
  • Advance enterprise resiliency through improved recovery capabilities
  • Reduce recovery time via automation
  • Enable rehoused recovery into new datacenters
  • Strengthen platform reliability through data protection design
  • Craft scalable solutions and automation to monitor the health and establish signals to drive understanding of Container Platform environments
  • Strengthen operational processes with client’s support and incident management teams for cloud ecosystem
  • Drive ongoing reliability improvements in their Kubernetes service offerings
  • Prioritize client / partner needs to serve as their voice and guide execution of the team
Required Qualifications
  • Bachelor’s Degree or equivalent experience in a technology related field (e.g. Computer Science, Engineering, etc.) required.
  • Production experience running Cloud and on‑prem Storage workloads at scale
  • Experience managing and maintaining Kubernetes Clusters on EKS/AKS and RKS.
  • Demonstrates a drive for continuous improvement and enjoys tackling complex problems.
  • Experience managing and interpreting large datasets using query languages and visualization tools (PowerBI/tableau)
  • Experience in software development with Python, NodeJS, or Java with a focus on SDLC and automation
  • 5 -7 years of hands‑on experience deploying and/or supporting highly distributed multi‑tiered systems at scale.
  • Experience building and deploying Docker images including Docker Compose
  • Hands‑on experience with Jenkins Core, including authoring and maintaining declarative CI/CD pipelines and libraries
  • Experience with distributed version control systems, Git preferred
  • Experience crafting and maintaining logging, monitoring, and alerting capabilities using tools like Datadog and Splunk
  • Practical experience in building cloud hosted and native applications for the enterprise.
  • Maintains a deep understanding of a wide variety of AWS/Azure services that support reliability, observability, and automation/orchestration.
  • Experience in incident/crisis management and supporting critically important applications
  • Hard on experience with one or more observability tools (Prometheus, Grafana, ELK/OpenSearch, OpenTelemetry, Datadog, etc.)
  • Ability to automate with various scripting languages (Python, Shell scripting, etc.)
  • Experience managing systems using infrastructure as code tools (IAM, ARM, Terraform, Chef)
Preferred Qualifications
  • Go, Angular, Python, JavaScript, AWS, RESTful services, Ruby, MVC, Jenkins CI/CD, Configuration Automation (Chef, Ansible)
  • Bootstrap, HTML/CSS, Shell Scripting, messaging frameworks (MQ), Service Oriented/Micro-service Architectures, OpenStack, Relational Databases (PostgreSQL)
  • Comfortable working in both Public and private cloud environments
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Site Reliability Engineer
Senior Site Reliability Engineer

O.C. Tanner • Salt Lake City (UT)

On-site
USD 130,000 - 180,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Jobtailor • Arlington (VA)

On-site
USD 140,000 - 200,000
Resiliency Architect
Resiliency Architect

ALLTECH CONSULTING SVC INC • Town of Texas (WI)

On-site
USD 100,000 - 130,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

ARA • Albuquerque (NM)

Hybrid
USD 120,000 - 180,000
Reliability Engineer (Remote)
Reliability Engineer (Remote)

Kohl's • Menomonee Falls (WI)

On-site
USD 95,000 - 135,000
Sr. Cloud Operations Reliability Engineer (SRE)
Sr. Cloud Operations Reliability Engineer (SRE)

NextGen Healthcare • Georgia

On-site
USD 140,000 - 210,000
Reliability Engineer
Reliability Engineer

Jobtailor • Reston (VA)

On-site
USD 140,000 - 190,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Veriipro • Charlotte (NC)

On-site
USD 140,000 - 190,000
Senior SRE (Contract/Hybrid)
Senior SRE (Contract/Hybrid)

Optomi • Orlando (FL)

Hybrid
USD 120,000 - 180,000
Senior Cloud Reliability Solutions Engineer
Senior Cloud Reliability Solutions Engineer

MathWorks • Natick (MA)

Hybrid
USD 140,000 - 200,000