Systems Reliability Engineer (Linux Admin)

The ICE Group

Plano (TX)

On-site

USD 80,000 - 110,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

The ICE Group in Plano, Texas is looking for a passionate Systems Reliability Engineer (SRE) to manage complex systems and maximize customer satisfaction. The role involves monitoring cluster health and ensuring service availability while being part of a 24/7 support rotation.

Ideal candidates will have at least 4 years of experience, particularly in distributed systems and databases. Strong scripting skills, preferably in Python, are essential for debugging and automation in this fast-paced environment.

Qualifications

  • Minimum 4 years of relevant experience.
  • Strong background in distributed systems and databases.
  • Excellent scripting skills for debugging and automation.

Responsibilities

  • Monitor cluster health and ensure availability of services.
  • Be on support rotation for 24/7 availability.
  • Contribute to knowledge bases and write tools/scripts for support.

Skills

Distributed systems
Performance analysis
Scripting (Python)
Server hardware troubleshooting

Job description

Systems Reliability Engineer (Linux Admin)
  • Full-time

We are seeking a well‑rounded Systems Reliability Engineer (SRE) who has a passion for working on complex systems. The right candidate for this position will have good communication as well as technical skills and will work hard to maximize customer satisfaction. The ability to work with cross‑functional teams in a rapidly growing environment is important.

The candidate will be responsible for monitoring cluster health and making sure services are always available. This person will have to be available on support 24/7 on a rotational basis. He/She will actively contribute to internal and external knowledge bases. Responsibilities sometimes require writing tools and scripts to solve customer issues and/or improve support infrastructure.

Skill Set
  • At least 4 years of relevant overall experience.
  • Background with distributed systems, databases and performance analysis.
  • Excellent scripting skills for debug and automation (Python knowledge is a plus).
  • Server hardware troubleshooting is a plus.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Linux SRE: 24/7 Reliability & Automation
Senior Linux SRE: 24/7 Reliability & Automation

The ICE Group • Plano (TX)

On-site
USD 80,000 - 110,000
Senior System Administrator / Site Reliability Engineer (SRE)
Senior System Administrator / Site Reliability Engineer (SRE)

Vitalwerks Internet Solutions, LLC • Reno (NV)

On-site
USD 80,000 - 110,000
Site Reliability Engineer
Site Reliability Engineer

SRE • Puerto Rico

Hybrid
USD 120,000 - 180,000
Linux System Admin with Python /SRE
Linux System Admin with Python /SRE

TechDigital Group • Town of Texas (WI)

Hybrid
USD 120,000 - 180,000
Site Reliability Engineer for Linux administration
Site Reliability Engineer for Linux administration

Hyve Solutions • Greenville (SC)

On-site
USD 60,000 - 90,000
Site Reliability Engineer for Linux administration
Site Reliability Engineer for Linux administration

Hyve Solutions • Clearwater (FL)

On-site
USD 70,000 - 90,000
Site Reliability Engineer
Site Reliability Engineer

TalentDome Staffing • United States

On-site
USD 140,000 - 210,000
Linux SRE: Systems Reliability & Automation
Linux SRE: Systems Reliability & Automation

Hyve Solutions • Greenville (SC)

On-site
USD 60,000 - 90,000
Site Reliability Engineer
Site Reliability Engineer

Brooksource • San Antonio (TX)

On-site
USD 80,000 - 120,000
Site Reliability Engineer
Site Reliability Engineer

System One • Pittsburgh

On-site
USD 140,000 - 190,000