Observability Engineer / Site Reliability Engineer

Jobtailor

California (MO)

On-site

USD 140,000 - 190,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Jobtailor is seeking an experienced SRE/Observability Engineer to architect, implement and maintain observability across cloud and hybrid environments. Your focus will be on Google Cloud Platform tooling (Cloud Logging, Cloud Monitoring, Trace and Profiler) and on building reliable dashboards with Grafana.

You will design and deploy robust observability stacks using Prometheus, Kubernetes (GKE/OpenShift) and cloud-native integrations, and drive IaC initiatives with Terraform and Ansible for

Qualifications

  • Proven experience with Google Cloud Platform and cloud-native monitoring.
  • Hands-on use of Grafana, Prometheus, and Google Cloud Observability suites.
  • Expertise in Linux/Unix administration and shell scripting.
  • Proficient in at least one modern language (Python/Go/Java).

Responsibilities

  • Architect, optimize, and maintain observability frameworks across cloud environments.
  • Design and deploy observability stacks across hybrid ecosystems.
  • Implement IaC with Terraform and Ansible for automated deployments.
  • Manage containerized apps on Kubernetes and GKE/OpenShift with CI/CD pipelines.
  • Define SLIs, SLOs, and error budgets to improve reliability.

Skills

GCP
Grafana
Prometheus
IaC
Kubernetes
Terraform
Ansible
Linux
Shell Scripting
Python

Tools

GKE
OpenShift
CI/CD
GitHub
Harness

Job description

  • Architect, optimize, and maintain observability frameworks across cloud environments, with a specific focus on implementing Google Cloud Platform (GCP) observability tools (Cloud Logging, Cloud Monitoring, Trace, and Profiler).
  • Design, deploy, and maintain robust observability stacks across hybrid ecosystems, utilizing Prometheus, Grafana, and cloud-native integrations.
  • Drive infrastructure-as-code (IaC) initiatives using Terraform and Ansible to ensure consistent, automated deployments of infrastructure and observability tooling.
  • Build, maintain, and optimize deployment workflows within Kubernetes and Google Kubernetes Engine (GKE) / OpenShift environments using GitHub, Harness, and other CI/CD pipelines.
  • Deeply analyze Linux/Unix system administration architectures, optimizing compute resource metrics and performance tuning across complex, distributed environments.
  • Implement SRE best practices, establishing meaningful Service Level Indicators (SLIs), Service Level Objectives (SLOs), and Error Budgets to ensure platform reliability.

Requirements

  • Proven engineering experience within Google Cloud Platform (GCP) environments, particularly managing cloud-native monitoring and compute resources.
  • Hands‑on experience with Grafana, Prometheus, and Google Cloud Observability suites.
  • Expert‑level knowledge of Linux/Unix operating systems paired with strong shell scripting skills for automation and system management.
  • Professional coding proficiency in at least one modern language (Python, Go, Java, Perl, or advanced Shell).
  • Hands‑on experience managing containerized applications on Kubernetes, GKE, and/or Red Hat OpenShift.

Core Competencies

Demonstrates expertise in architecting and optimizing observability frameworks within Google Cloud Platform (GCP) environments, utilizing tools such as Grafana and Prometheus. Proficient in implementing infrastructure-as-code (IaC) practices with Terraform and Ansible, and managing containerized applications in Kubernetes and GKE.

Highest-signal resume keywords

  • Google Cloud Platform (GCP)
  • Grafana
  • Prometheus
  • Infrastructure-as-Code (IaC)
  • Kubernetes

ATS Optimization Keywords

Hard Skills

  • Linux/Unix System Administration
  • Shell Scripting
  • Python
  • Go
  • Java
  • Perl
  • Terraform
  • Ansible
  • Cloud Logging
  • Cloud Monitoring

Industry Keywords

  • Observability Frameworks
  • Service Level Indicators (SLIs)
  • Service Level Objectives (SLOs)
  • Error Budgets
  • Cloud-Native Integrations

Tools & Technologies

  • Google Kubernetes Engine (GKE)
  • OpenShift
  • CI/CD Pipelines
  • GitHub
  • Harness
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Observability Engineer / Site Reliability Engineer
Observability Engineer / Site Reliability Engineer

Ontrac Solutions • Chicago (IL)

On-site
USD 120,000 - 180,000
Senior DevOps Engineer, AI Platform
Senior DevOps Engineer, AI Platform

Jobtailor • Houston (TX)

On-site
USD 150,000 - 210,000
Staff Software Engineer, Platform & Solutions
Staff Software Engineer, Platform & Solutions

twentysix • El Segundo (CA)

On-site
USD 180,000 - 240,000
IT Manager – Platform Engineering, Data Science
IT Manager – Platform Engineering, Data Science

Jobtailor • Newport News (VA)

On-site
USD 180,000 - 240,000
Observability Engineer
Observability Engineer

Gardner Resources Consulting, LLC • Massachusetts

On-site
Staff Site Reliability Engineer
Staff Site Reliability Engineer

Jobtailor • California (MO)

On-site
USD 140,000 - 210,000
Linux Team Lead
Linux Team Lead

Jobtailor • Glen Burnie (MD)

On-site
USD 120,000 - 150,000
Senior Site Reliability Engineer NEX
Senior Site Reliability Engineer NEX

Patterson-UTI • Houston (TX)

On-site
USD 120,000 - 180,000
Site Reliability Engineer
Site Reliability Engineer

Jobtailor • United States

On-site
USD 120,000 - 180,000
Senior Platform Engineer – Core Infrastructure
Senior Platform Engineer – Core Infrastructure

Jobtailor • California (MO)

On-site
USD 150,000 - 190,000