Site Reliability Engineer – Observability & Automation Lead

Talentify

Houston (TX)

On-site

USD 140,000 - 180,000

Full time

14 days+
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Job summary

Talentify is seeking a Site Reliability Engineer in Houston to own reliability, observability, and performance across distributed systems. You will lead incident response, design health checks, and drive postmortems while collaborating with developers to evolve architecture for scale.

The role requires 5+ years in SRE/DevOps, strong Kubernetes and AKKA.NET experience, and hands-on cloud and PostgreSQL skills.

Qualifications

  • 5+ years of experience in SRE, DevOps, or Infrastructure Engineering roles.
  • Expertise in Kubernetes and container orchestration at scale.
  • Strong experience with AKKA.NET or similar actor-based frameworks.
  • Proficiency with scripting and automation (Bash, PowerShell, Python).
  • Experience with observability tools (Phobos, Datadog, Prometheus, Grafana, OpenTelemetry, ELK).
  • Hands-on experience with cloud platforms (AWS, Azure, or GCP).
  • Strong PostgreSQL knowledge—performance tuning, query optimization, maintenance.

Responsibilities

  • Maintain and monitor production systems for availability, latency, and performance.
  • Lead incident response efforts, including communication, resolution, and postmortem documentation.
  • Design and implement health checks, alerting systems, and automated remediation workflows.
  • Drive root cause analysis and implement permanent resolutions for recurring issues.
  • Set up and maintain full observability stacks (logging, metrics, tracing).
  • Tune and optimize distributed systems for performance and resource efficiency.
  • Collaborate with developers to evolve architecture and improve throughput, latency, and stability.
  • Design and maintain modern CI/CD pipelines and automate deployment, testing, rollback processes.

Skills

SRE experience
Kubernetes
AKKA.NET
Bash
PowerShell
Python
Datadog
Prometheus
Grafana
OpenTelemetry
ELK
AWS
Azure
GCP
PostgreSQL
Incident management
Ownership

Tools

GitHub Actions
Azure Pipelines
GitLab CI
Terraform
Docker

Job description

Talentify is seeking a Site Reliability Engineer in Houston to own reliability, observability, and performance across distributed systems. You will lead incident response, design health checks, and drive postmortems while collaborating with developers to evolve architecture for scale.

The role requires 5+ years in SRE/DevOps, strong Kubernetes and AKKA.NET experience, and hands-on cloud and PostgreSQL skills.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Site Reliability Engineer
Site Reliability Engineer

Talentify • Houston (TX)

On-site
USD 140,000 - 180,000
Site Reliability Engineer: Cloud, Automation & Observability
Site Reliability Engineer: Cloud, Automation & Observability

TechDigital Group • Houston (TX), Juno Beach (FL)

On-site
USD 120,000 - 180,000
Site Reliability Engineer: Distributed Systems Observability
Site Reliability Engineer: Distributed Systems Observability

engineeringjobs.net, Inc. • Sunnyvale (CA)

On-site
USD 140,000 - 210,000
Senior Site Reliability Engineer – Scale & Observability
Senior Site Reliability Engineer – Scale & Observability

Inspire Brands, Inc. • Atlanta (GA)

On-site
USD 120,000 - 180,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Optomi • Dallas (TX)

On-site
USD 120,000 - 150,000
Site Reliability Engineer
Site Reliability Engineer

NextGen | GTA: A Kelly Telecom Company • Mount Laurel Township (NJ)

On-site
USD 110,000 - 170,000
Remote Site Reliability Engineer - Observability Expert
Remote Site Reliability Engineer - Observability Expert

Codevertex Innovations • Northern (KY)

Hybrid
USD 140,000 - 195,000
Senior Site Reliability Engineer: Scale, Automate, Observe
Senior Site Reliability Engineer: Scale, Automate, Observe

TalentDome Staffing • United States

On-site
USD 140,000 - 210,000
Senior Site Reliability Engineer – Observability & Cloud
Senior Site Reliability Engineer – Observability & Cloud

Cosm Inc. • El Segundo (CA), Northern (KY)

Hybrid
USD 110,000 - 145,000
Sr. Site Reliability Engineer
Sr. Site Reliability Engineer

MeridianLink • United States

On-site
USD 140,000 - 190,000