Observability Engineer Site Reliability Engineer

Ontrac Solutions

Arizona

Hybrid

USD 140,000 - 190,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Verification cost reimbursement

Job summary

Ontrac Solutions is seeking an experienced Observability / Site Reliability Engineer (SRE) to design, scale, and maintain enterprise monitoring and alerting ecosystems across multi-cloud and native systems.

You will bridge development and operations, ensuring high availability, performance tuning, and deep visibility while driving automation and robust observability pipelines using cloud-native tools. The role emphasizes SRE best practices and measurable reliability targets.

Qualifications

  • Experience designing and operating cloud-based observability platforms.
  • Strong Linux/Unix administration and scripting skills.
  • Proficiency in Terraform and Ansible for IaC.
  • Hands-on with Kubernetes, GKE or OpenShift.
  • Proficiency with Prometheus and Grafana, plus Google Cloud Observability suites.
  • Programming in Python/Go/Java.

Responsibilities

  • Architect and maintain observability stacks across hybrid/multi-cloud environments.
  • Define SLIs, SLOs, and error budgets to ensure reliability.
  • Automate deployments and scaling with Terraform and Ansible.
  • Integrate CI/CD pipelines in Kubernetes/GKE/OpenShift.
  • Analyze and optimize system performance and resource utilization.
  • Collaborate with development and operations to improve reliability.

Skills

Linux/Unix mastery
Shell scripting
Programming (Python/Go/Java)
SRE practices
CI/CD concepts

Education

Bachelor's degree in CS/Engineering

Tools

Prometheus
Grafana
Google Cloud Monitoring
Cloud Logging
Cloud Tracing
Cloud Profiler
Kubernetes
GKE/OpenShift
Terraform
Ansible

Job description

About Ontrac Solutions

Ontrac Solutions is a leading technology consulting firm, specializing in cutting-edge solutions that drive business transformation. We partner with organizations to modernize their infrastructure, streamline processes, and deliver tangible results. By creating value beyond the hype, we help businesses modernize technology and build new strategies that fuel growth. Our team is committed to innovation, collaboration, and excellence, empowering our clients to succeed in an evolving digital landscape.

Role Overview

We are seeking an experienced Observability / Site Reliability Engineer (SRE) to design, scale, and maintain our enterprise monitoring and alerting ecosystems. In this role, you will bridge the gap between development and operations by ensuring high availability, performance tuning, and deep visibility across distributed multi-cloud and native systems. You will play a critical role in automating infrastructure and building robust observability pipelines using industry-leading cloud-native tools.

Key Responsibilities
  • GCP & Cloud Management: Architect, optimize, and maintain observability frameworks across cloud environments, with a specific focus on implementing Google Cloud Platform (GCP) observability tools (Cloud Logging, Cloud Monitoring, Trace, and Profiler).
  • Platform Management: Design, deploy, and maintain robust observability stacks across hybrid ecosystems, utilizing Prometheus, Grafana, and cloud-native integrations.
  • Automation & IaC: Drive infrastructure-as-code (IaC) initiatives using Terraform and Ansible to ensure consistent, automated deployments of infrastructure and observability tooling.li>
  • CI/CD Integration: Build, maintain, and optimize deployment workflows within Kubernetes and Google Kubernetes Engine (GKE) / OpenShift environments using GitHub, Harness, and other CI/CD pipelines.
  • System Performance: Deeply analyze Linux/Unix system administration architectures, optimizing compute resource metrics and performance tuning across complex, distributed environments.
  • SRE Evangelism: Implement SRE best practices, establishing meaningful Service Level Indicators (SLIs), Service Level Objectives (SLOs), and Error Budgets to ensure platform reliability.
Required Skills & Qualifications
  • Cloud Infrastructure: Proven engineering experience within Google Cloud Platform (GCP) environments, particularly managing cloud-native monitoring and compute resources.
  • Observability Tooling: Hands-on experience with Grafana, Prometheus, and Google Cloud Observability suites. Direct experience with GEM (Grafana Enterprise Metrics) is highly desirable.
  • OS & Scripting: Expert-level knowledge of Linux/Unix operating systems paired with strong shell scripting skills for automation and systems management.
  • Programming: Professional coding proficiency in at least one modern language (Python, Go, Java, Perl, or advanced Shell).
  • Containers & Orchestration: Hands-on experience managing containerized applications on Kubernetes, GKE, and/or Red Hat OpenShift.

Ontrac Solutions has partnered with PinpointVerify to help genuine applicants rise above the noise. Today, qualified candidates are too often overlooked due to fake and fraudulent applications. PinpointVerify gives recruiters confidence that you are exactly who you say you are — and gives you a portable verification credential you can share with any employer.

Applicants who complete verification are prioritized over non-verified candidates with comparable experience. And if you're hired, Ontrac reimburses the full cost of your verification.
Get verified https://pinpointverify.com/ontrac

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Observability Engineer Site Reliability Engineer
Observability Engineer Site Reliability Engineer

Ontrac Solutions • United States

On-site
USD 120,000 - 180,000
Observability Engineer / Site Reliability Engineer
Observability Engineer / Site Reliability Engineer

Ontrac Solutions • Chicago (IL)

On-site
USD 120,000 - 180,000
Observability & SRE Engineer — Cloud, Kubernetes, & Automation
Observability & SRE Engineer — Cloud, Kubernetes, & Automation

Ontrac Solutions • Arizona

Hybrid
USD 140,000 - 190,000
Verification cost reimbursement
SRE/Observability Engineer
SRE/Observability Engineer

BlueSky Resource Solutions • United States

Remote
USD 100,000 - 130,000
Senior Observability & SRE Engineer — GCP/Kubernetes
Senior Observability & SRE Engineer — GCP/Kubernetes

Ontrac Solutions • United States

On-site
USD 120,000 - 180,000
Senior Cloud Observability & SRE Engineer
Senior Cloud Observability & SRE Engineer

Ontrac Solutions • Chicago (IL)

On-site
USD 120,000 - 180,000
Associate Engineer, Site Reliability
Associate Engineer, Site Reliability

R&D • United States

On-site
USD 90,000 - 140,000
Staff Site Reliability Engineer - Observability GCP
Staff Site Reliability Engineer - Observability GCP

Okta • Bellevue (WA)

On-site
USD 174,000 - 239,000
Health insurance
401(k)
Paid leave
+1
Site Reliability Engineer (SRE) - Observability Specialist at Vodastra Las Vegas, NV
Site Reliability Engineer (SRE) - Observability Specialist at Vodastra Las Vegas, NV

Downtown Boulder Partnership • Las Vegas (NV)

On-site
Senior Site Reliability Engineer, Observability
Senior Site Reliability Engineer, Observability

blockchaincapital.com • New York (NY)

On-site
USD 130,000 - 180,000