Remote SRE — Incident Response, Reliability & Observability

Pyramid Consulting, Inc

United States

Remote

MXN 420,000 - 640,000

Full time

36 hours ago
Be an early applicant
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Job summary

Pyramid Consulting, Inc. is seeking a Site Reliability Engineer (SRE) to improve the stability and performance of large-scale production systems.

This remote role focuses on incident response, automation, observability, and data analysis to drive reliability improvements across cloud-native environments. The role requires strong Python scripting, advanced SQL, and experience with RCA, monitoring, logging, and distributed tracing.

Qualifications

  • 4+ years of experience in SRE, DevOps, Systems Engineering, or Production Support.
  • Strong Python scripting and automation experience.
  • Advanced SQL and data analysis skills.
  • Experience with incident response, RCA, and production troubleshooting.
  • Strong understanding of observability, monitoring, logging, and distributed tracing.
  • Experience with Kubernetes and Google Cloud Platform (GCP).
  • Experience integrating and working with APIs.

Responsibilities

  • Respond to and manage production incidents, perform root cause analysis, and drive preventative solutions.
  • Develop automation and operational tooling using Python and APIs.
  • Analyze incident, log, and performance data using SQL and statistical methods to identify reliability improvements.
  • Build and enhance monitoring, alerting, logging, and distributed tracing capabilities.
  • Support and optimize Kubernetes-based applications running in GCP.
  • Create dashboards, reports, and reliability metrics to communicate technical and business impact.
  • Partner with engineering teams to improve system resilience, scalability, and operational efficiency.

Skills

Python
SQL
Incident response
RCA
APIs
Kubernetes
GCP
Observability
Monitoring
Logging
Distributed tracing
SRE
Production Support

Tools

Datadog
Splunk
Grafana
Prometheus
OpenTelemetry

Job description

Pyramid Consulting, Inc. is seeking a Site Reliability Engineer (SRE) to improve the stability and performance of large-scale production systems.

This remote role focuses on incident response, automation, observability, and data analysis to drive reliability improvements across cloud-native environments. The role requires strong Python scripting, advanced SQL, and experience with RCA, monitoring, logging, and distributed tracing.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Site Reliability Engineer
Senior Site Reliability Engineer

Pyramid Consulting, Inc • United States

Remote
MXN 420,000 - 640,000
SRE - Incident RCA & Observability (Remote)
SRE - Incident RCA & Observability (Remote)

Socket.dev • Georgia

Hybrid
USD 100,000 - 120,000
Medical and Dental coverage
Vision coverage
401(k) match
+8
Remote SRE — Scale, Resilience & Observability
Remote SRE — Scale, Resilience & Observability

Bright Vision Technologies • United States

On-site
USD 100,000 - 150,000
Competitive base salary
Health benefits
Long-term stability
Remote Senior SRE: Reliability & Observability
Remote Senior SRE: Reliability & Observability

United States Digital Space LLC • United States

Remote
USD 120,000 - 180,000
Remote-first
Hybrid work options
EU time-zone considerations
Remote SRE: Production Systems & AI Observability
Remote SRE: Production Systems & AI Observability

Crossing Hurdles • United States

Remote
USD 80,000 - 100,000
Site Reliability Engineer (SRE)
Site Reliability Engineer (SRE)

ACI Infotech • Seattle (WA), Northern (KY)

Hybrid
USD 120,000 - 170,000
Health insurance
Dental insurance
Vision insurance
+3
SRE Lead: Production Reliability & Observability Architect
SRE Lead: Production Reliability & Observability Architect

TechDigital Group • Woonsocket (RI)

On-site
USD 140,000 - 190,000
Remote SRE Manager: Lead Incidents & Reliability
Remote SRE Manager: Lead Incidents & Reliability

NationsBenefits India • United States

On-site
USD 150,000 - 210,000
Unlimited PTO
Fully remote for US-based employees
Competitive compensation
+1
Hybrid SRE: Infra & Assurance Services
Hybrid SRE: Infra & Assurance Services

TikTok • Seattle (WA)

Hybrid
USD 112,000 - 178,000
Remote SRE — Cloud Reliability & Performance
Remote SRE — Cloud Reliability & Performance

DevOpsChat • United States

Hybrid
USD 120,000 - 170,000
Healthcare options
Professional development
Flexible work location
+1