Site Reliability Engineer

Ethos Group

Irving (TX)

On-site

USD 110,000 - 160,000

Full time

36 hours ago
Be an early applicant
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Job summary

Ethos Group is seeking a Site Reliability Engineer (SRE) to join our growing technology team in Irving, TX. You will focus on building highly available, scalable systems and partner with development and infrastructure teams to improve performance and automation.

In this role you will maintain uptime, optimize applications, automate operations, and support a modern cloud-based environment, contributing to operational excellence and reliability across our platforms.

Qualifications

  • Bachelor's degree or equivalent in CS/IT.
  • Experience in SRE, DevOps, or cloud infra roles.
  • Strong Linux and cloud tech understanding.
  • Hands-on Kubernetes, CI/CD, and IaC experience.

Responsibilities

  • Design and maintain scalable cloud infrastructure.
  • Monitor performance and ensure high availability.
  • Automate deployments and runbooks.
  • Participate in on-call and incident response.
  • Improve observability and incident analysis.

Skills

Cloud infra
Linux
SRE/DevOps
Python
Communication
Troubleshooting
On-call

Education

Bachelor's degree in CS/IT

Tools

Terraform
CloudFormation
Kubernetes
Azure Resource Manager
CI/CD tools

Job description

Ethos Group is seeking a talented and proactive Site Reliability Engineer (SRE) to join our growing technology team.

This role is ideal for an engineer who is passionate about building highly available, scalable, and reliable systems while partnering closely with development and infrastructure teams to improve platform performance, automation, and operational excellence.

As a Site Reliability Engineer, you will play a critical role in maintaining system uptime, optimizing application performance, automating operational processes, and supporting a modern cloud-based technology environment.

Key Responsibilities
  • Design, implement, and maintain reliable, scalable, and secure cloud infrastructure
  • Monitor application and system performance to ensure maximum availability
  • Automate operational tasks and deployment processes
  • Respond to incidents, troubleshoot production issues, and drive root cause analysis
  • Develop and maintain monitoring, alerting, and observability solutions
  • Partner with software engineering teams to improve application reliability and performance
  • Support CI/CD pipelines and deployment automation initiatives
  • Create operational documentation, runbooks, and best practices
  • Participate in on-call rotations and incident response activities
  • Continuously identify opportunities for system improvements and increased efficiency
Qualifications
  • Bachelor's degree in Computer Science, Information Technology, or related field, or equivalent experience
  • Experience in Site Reliability Engineering, DevOps, Systems Engineering, or Cloud Infrastructure roles
  • Strong understanding of Linux and cloud technologies
  • Direct hands-on experience with AWS, Azure, or Google Cloud Platform
  • Experience writing and maintaining infrastructure-as-code tools such as Azure resource templates, Terraform, or CloudFormation
  • Direct hands-on experience supporting CI/CD pipelines and automation tools
  • Direct hands-on experience with production grade Kubernetes deployments
  • Experience with monitoring and observability platforms
  • Strong scripting skills in Python, PowerShell, Bash, or similar languages
  • Excellent troubleshooting and problem-solving abilities
  • Strong communication and collaboration skills
Preferred Qualifications
  • Deep knowledge of networking, security, and cloud architecture principles
  • Experience with incident management and root cause analysis
  • Relevant cloud, Kubernetes, or DevOps certifications
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Site Reliability Engineer
Site Reliability Engineer

TalentDome Staffing • United States

On-site
USD 140,000 - 210,000
Site Reliability Engineer (SRE)
Site Reliability Engineer (SRE)

myBridge Corporation • Austin (TX)

On-site
USD 120,000 - 160,000
Site Reliability Engineer
Site Reliability Engineer

NextGen | GTA: A Kelly Telecom Company • Mount Laurel Township (NJ)

On-site
USD 110,000 - 170,000
Site Reliability Engineer
Site Reliability Engineer

Harrison Clarke • New York (NY)

On-site
USD 120,000 - 160,000
Site Reliability Engineering (SRE)
Site Reliability Engineering (SRE)

Weekday (YC W21) • New York (NY)

On-site
USD 150,000 - 250,000
Health, dental, vision insurance
Generous PTO
Learning & development
+2
Site Reliability Engineer
Site Reliability Engineer

Brooksource • San Antonio (TX)

On-site
USD 80,000 - 120,000
Senior SRE Engineer
Senior SRE Engineer

Compunnel, Inc. • Alpharetta (GA)

On-site
USD 140,000 - 190,000
Site Reliability Engineer
Site Reliability Engineer

Evlo AI • Minneapolis (MN)

On-site
USD 120,000 - 180,000
Sr SRE Automation Engineer
Sr SRE Automation Engineer

Compunnel, Inc. • Austin (TX), Northern (KY)

Hybrid
USD 130,000 - 180,000
Sr. Site Reliability Engineer
Sr. Site Reliability Engineer

Mike Albert Fleet Solutions • Cincinnati (OH)

On-site
USD 100,000 - 135,000