Senior Site Reliability Engineer NEX

Patterson-UTI

Houston (TX)

On-site

USD 120,000 - 180,000

Full time

7 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Patterson-UTI is seeking a Reliability Engineer to design, implement, and operate scalable, resilient systems on Google Cloud Platform in a 24/7 production environment. You will drive SRE practices, define SLOs/SLIs, and own incident response and postmortems while collaborating with software, security, and product teams.

You will also contribute to capacity planning, CI/CD improvements, and infrastructure automation using Terraform and Kubernetes, ensuring secure access and efficient operations

Qualifications

  • 3+ years in Site Reliability Engineering, platform engineering, DevOps, or similar.

Responsibilities

  • Design, implement, and operate scalable, resilient systems on Google Cloud Platform.
  • Improve service availability, latency, performance, scalability, and operational resilience.
  • Define, implement, and track service-level indicators, service-level objectives, and error budgets.
  • Perform capacity planning, performance analysis, and workload forecasting.
  • Design and validate disaster recovery, backup, failover, and service-restoration capabilities.
  • Implement and maintain secure cloud networking, IAM, workload identities, service accounts, and access-control practices.
  • Partner with cybersecurity and identity teams to ensure infrastructure and services follow organizational security standards.
  • Monitor cloud consumption and optimize resource utilization, performance, and cost efficiency.
  • Identify operational risks and recommend improvements to cloud architecture and service design.

Skills

SRE/DevOps
Incident management
On-call experience
English proficiency
Stakeholder communication

Education

Bachelor's degree in Computer Science or related field

Tools

Terraform
Docker
Kubernetes
GitHub Actions
Azure DevOps
Bitbucket Pipelines
Python
Go
Java

Job description

Reliability Engineering
  • Design, implement, and operate scalable, resilient, and highly available systems on Google Cloud Platform.
  • Improve service availability, latency, performance, scalability, and operational resilience.
  • Define, implement, and track service-level indicators, service-level objectives, and error budgets.
  • Perform capacity planning, performance analysis, and workload forecasting.
  • Design and validate disaster recovery, backup, failover, and service-restoration capabilities.
  • Implement and maintain secure cloud networking, IAM, workload identities, service accounts, and access-control practices.
  • Partner with cybersecurity and identity teams to ensure infrastructure and services follow organizational security standards.
  • Monitor cloud consumption and optimize resource utilization, performance, and cost efficiency.
  • Identify operational risks and recommend improvements to cloud architecture and service design.
Automation and Platform Engineering
  • Build and maintain cloud infrastructure using Terraform or comparable infrastructure-as-code tools.
  • Automate repetitive operational activities and systematically identify, measure, and reduce manual toil.
  • Build reusable infrastructure modules, deployment patterns, and operational tooling.
  • Improve CI/CD pipelines to enable secure, repeatable, and reliable software delivery.
Observability and Incident Management
  • Develop actionable alerts that identify meaningful service degradation while reducing alert fatigue and unnecessary operational noise.
  • Create and maintain dashboards, runbooks, operational procedures, and troubleshooting documentation.
  • Participate in a sustainable on-call rotation supporting production systems.
  • Respond to production incidents, coordinate service restoration, and lead incident response when appropriate.
  • Facilitate blameless postmortems and identify corrective and preventive actions.
  • Use incident and operational data to improve system design, automation, monitoring, and response processes.
Collaboration and Service Ownership
  • Partner with software engineering, data engineering, security, and product teams to improve application reliability and production operations.
  • Promote shared responsibility for production reliability between application development and platform teams.
  • Establish and document reliability standards, operational practices, and reusable engineering patterns.
  • Provide technical guidance and coaching on SRE, cloud, Kubernetes, observability, and incident-management practices.
Required Knowledge, Skills, and Abilities
  • Three or more years of experience in Site Reliability Engineering, platform engineering, DevOps, cloud engineering, production software engineering, or a similar role.
  • Experience operating highly available systems in a 24/7 production environment.
  • Hands‑on experience operating workloads on Google Cloud Platform or another major public cloud platform.
  • Strong experience managing compute, networking and data GCP services workloads
  • Strong experience with containerization and orchestration technologies, including Docker and Kubernetes.
  • Experience building and managing infrastructure with Terraform or a comparable infrastructure-as-code tool.
  • Proficiency in Python, Go, Java, or another comparable programming language.
  • Experience implementing or operating CI/CD pipelines using GitHub Actions,Azure DevOps, Bitbucket Pipelines, or comparable tools.
  • Experience implementing observability using metrics, logs, traces, dashboards, and alerts.
  • Experience participating in on‑call rotations, responding to incidents, and contributing to postmortems.
  • Understanding of SLIs, SLOs, error budgets, and other SRE principles.
  • Ability to troubleshoot complex issues across application, infrastructure, network, data, and cloud‑service layers.
  • Ability to communicate effectively with engineering teams, business stakeholders, and operational personnel.
Minimum Qualifications
  • Bachelor's degree in Computer Science, Information Technology, Engineering, or a related field, or equivalent practical experience.
  • 3+ years of experience in Site Reliability Engineering, platform engineering, cloud engineering, or DevOps.
  • 3+ years of experience operating production workloads in GCP.
  • Ability to understand and communicate in English at a level sufficient to issue, receive, and respond to safety‑related and operations‑related instructions.
Preferred Qualifications
  • Google Cloud and/or Kubernetes certifications.
  • Experience supporting data‑intensive, streaming, analytics, or event‑driven platforms.
  • Experience establishing production‑readiness, incident‑management, or reliability‑review processes.
  • Experience working in the energy, oil and gas, industrial, IoT, field operations, or other operationally critical industries.
  • Experience supporting technology environments that integrate cloud platforms with remote sites, field equipment, industrial systems, or edge computing.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Sr. Cloud Operations Reliability Engineer (SRE)
Sr. Cloud Operations Reliability Engineer (SRE)

NextGen Healthcare • Georgia

On-site
USD 140,000 - 210,000
Site Reliability Engineer
Site Reliability Engineer

Compunnel, Inc. • New Jersey

On-site
USD 120,000 - 150,000
Manager of Site Reliability Engineering (SRE)
Manager of Site Reliability Engineering (SRE)

Genuine Parts Company • Alabama

On-site
USD 120,000 - 150,000
Site Reliability Engineer
Site Reliability Engineer

JPS Tech Solutions • San Jose (CA)

On-site
USD 130,000 - 160,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Veloc Inc • Coppell (TX)

On-site
USD 140,000 - 190,000
Lead Site Reliability Engineer - Infrastructure & DevOps
Lead Site Reliability Engineer - Infrastructure & DevOps

SRI Tech Solutions Inc. • Orlando (FL)

On-site
USD 140,000 - 190,000
Senior Site Reliability Engineer (GCP Cloud)
Senior Site Reliability Engineer (GCP Cloud)

United States Digital Space LLC • United States

Hybrid
USD 120,000 - 180,000
Site Reliability Engineer
Site Reliability Engineer

TalentDome Staffing • United States

On-site
USD 140,000 - 210,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Bank of America • Charlotte (NC)

On-site
USD 152,000 - 192,000
Industry-leading benefits
Paid time off
Discretionary incentive eligibility
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Hobbsnews • Jersey City (NJ)

On-site
USD 152,000 - 192,000
Industry-leading benefits
Access to paid time off
Annual discretionary incentives