Site Reliability Engineer

INSPYR Solutions

Houston (TX)

Hybrid

USD 83,000 - 131,000

Full time

27 hours ago
Be an early applicant
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Benefits offered by this job

Comprehensive medical benefits
Competitive pay
401(k) retirement plan

Job summary

INSPYR Solutions is seeking a Site Reliability Engineer in Houston, TX for a hybrid contract-to-hire role. You will join a founding SRE team, moving from reactive operations toward engineered reliability, with on-call and incident response responsibilities.

You will own SLIs/SLOs, observability, and automation across cloud and on-prem environments, collaborating with application, infra, and operations teams to reduce incidents over time.

Qualifications

  • 4+ years of SRE/DevOps or related roles.
  • Strong Linux/Windows administration and troubleshooting.
  • Experience with automation and scripting.

Responsibilities

  • Define and implement SLI/SLOs with meaningful health metrics.
  • Ensure observability across cloud, on-prem, and APIs.
  • Automate toil: patching, data corrections, restarts, recovery tasks.
  • Develop self-healing behaviors for common failures.
  • Participate in on-call rotations and post-incident reviews.
  • Design and execute disaster recovery tests across environments.

Skills

SRE
DevOps
Automation
Observability
Incident response
On-call

Tools

Terraform
Ansible
Kubernetes

Job description

Title:Site Reliability Engineer

Location: Houston, TX 77002 (Hybrid: 3 days onsite / 2 days remote)

Duration: Contract to Hire

Work Requirements:U.S.Citizen, GC Holders,or Authorized to Work in the U.S.

Job Description

The Site Reliability Engineer is a founding member of SRE practice's.
This role exists to move the organization from reactive operations to engineered reliability. You will study how our most critical systems fail, particularly our internal applications and facility automation interfaces and design controls, automation, and observability that reduce incidents over time.
Success in this role means fewer false alerts, faster recovery, less manual intervention, and systems that heal themselves when possible.
You will work closely with application, infrastructure, and operations teams and participate directly in on call and incident response.

What You Will Own
  • Definition and implementation of SLIs and SLOs that measure meaningful system health, not just availability
  • Observability across the full stack, correlating cloud services, APIs, and on premise facility operations
  • Automation to eliminate operational toil, including patching, data corrections, restarts, and recovery tasks
  • Development of self healing behaviors for common failure modes
  • Participation in on call rotations and leadership of blameless post incident reviews
  • Design and execution of disaster recovery tests across SaaS, cloud, and on premise environments

This is hands on reliability engineering. The systems you improve will directly impact daily warehouse operations.

Technical Environment
  • Hybrid environments spanning cloud and on-premise infrastructure
  • Azure cloud services
  • Software Development/OOP skills within either .NET/Java/Python
  • Observability tooling across logs, metrics, and alerting
  • Automation using Python, PowerShell, Bash, or Ansible
  • CI/CD tools and modern deployment practices
  • Exposure to containerized and distributed systems environments
What We're Looking For
  • 4+ years of experience in SRE, DevOps, Systems Engineering, or related roles
  • Strong Linux and Windows systems administration and troubleshooting skills
  • Hands-on experience with automation and scripting
  • Experience designing and operating monitoring, alerting, and observability solutions
  • Practical experience working in Azure environments
  • Strong analytical skills and a bias toward eliminating root causes, not symptoms
  • Ability to collaborate across application, infrastructure, and operations teams
  • Exposure to Kubernetes, microservices, or container orchestration
  • Hands-on experience with infrastructure as code tools such as Terraform or Ansible
  • Understanding of distributed systems and high availability design
  • Experience with SRE practices such as SLO based operations, runbook automation, or chaos testing
Our benefits package includes
  • Comprehensive medical benefits
  • Competitive pay
  • 401(k) retirement plan
  • …and much more!
About INSPYR Solutions

Technology is our focus and quality is our commitment. As a national expert in delivering flexible technology and talent solutions, we strategically align industry and technical expertise with our clients' business objectives and cultural needs. Our solutions are tailored to each client and include a wide variety of professional services, project, and talent solutions. By always striving for excellence and focusing on the human aspect of our business, we work seamlessly with our talent and clients to match the right solutions to the right opportunities. Learn more about us at inspyrsolutions.com.

INSPYR Solutions provides Equal Employment Opportunities (EEO) to all employees and applicants for employment without regard to race, color, religion, sex, national origin, age, disability, genetics, or any other protected status. INSPYR Solutions complies with all applicable laws governing nondiscrimination in employment in every location in which the company has facilities.

Applicants requiring reasonable accommodation during the application or interview process should contact HR@inspyrsolutions.com for assistance.

Information collected and processed through your application with INSPYR Solutions (including any job applications you choose to submit) is subject to INSPYR Solutions' Privacy Policy and INSPYR Solutions' AI and Automated Employment Decision Tool Policy: https://www.inspyrsolutions.com/policies/. By submitting an application, you are consenting to being contacted by INSPYR Solutions through phone, email, or text.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

SRE/Devops Engineer
SRE/Devops Engineer

INSPYR Solutions • Sunnyvale (CA)

Hybrid
USD 120,000 - 180,000
Work-life balance
No on-call requirements
Standard business hours
Server Operations Specialist III
Server Operations Specialist III

INSPYR Solutions • Houston (TX)

On-site
USD 80,000 - 100,000
Comprehensive medical benefits
Competitive pay
401(k)
+1
Cloud Engineer
Cloud Engineer

INSPYR Solutions • Houston (TX)

Hybrid
USD 124,000 - 193,000
Comprehensive medical benefits
Competitive pay
401(k) Retirement plan
Site Reliability Engineer (SRE)
Site Reliability Engineer (SRE)

Randstad Digital Americas • Plano (TX)

On-site
USD 115,000 - 125,000
Medical insurance
401K plan
Lead Site Reliability Engineer (SRE) / Principal Site Reliability Engineer (SRE)
Lead Site Reliability Engineer (SRE) / Principal Site Reliability Engineer (SRE)

Mindlance • Irving (TX)

Hybrid
USD 120,000 - 160,000
Site Reliability Engineer (SRE)
Site Reliability Engineer (SRE)

Bright Vision Technologies • Cranberry Township

Remote
USD 100,000 - 180,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Spectraforce Technologies • Austin (TX)

Hybrid
USD 130,000 - 170,000
Infrastructure Project Manager (Disaster Recovery)
Infrastructure Project Manager (Disaster Recovery)

INSPYR Solutions • Houston (TX)

On-site
USD 110,000 - 140,000
Comprehensive medical benefits
401(k)
Retirement plan
Hybrid Site Reliability Engineer - Build Self-Healing
Hybrid Site Reliability Engineer - Build Self-Healing

INSPYR Solutions • Houston (TX)

Hybrid
USD 83,000 - 131,000
Comprehensive medical benefits
Competitive pay
401(k) retirement plan
Cloud Infrastructure Architect/Engineer – Multi-Cloud (Azure & AWS)
Cloud Infrastructure Architect/Engineer – Multi-Cloud (Azure & AWS)

INSPYR Solutions • Houston (TX)

On-site
USD 120,000 - 150,000
Comprehensive medical benefits
Competitive pay, 401(k)
Retirement plan