Site Reliability Engineer (SRE)

Smart IMS Inc

Southlake (TX)

On-site

USD 55,000 - 110,000

Full time

14 days+
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Job summary

Smart IMS Inc. in Southlake, TX is seeking a Site Reliability Engineer (SRE) for a 12-month contract to improve reliability, scalability, and efficiency of production systems through automation and observability.

The ideal candidate has strong Python development skills, hands-on production engineering experience, and a proven track record of reducing toil through automation. Onsite role with competitive hourly pay.

Qualifications

  • Bachelor's degree in Computer Science, Engineering, or a related field, or equivalent professional experience.
  • 3-5 years of hands-on Site Reliability Engineering or Production Engineering experience supporting large-scale production systems.
  • Strong Python programming expertise with demonstrated experience building automation tools, frameworks, and operational solutions.
  • Proven experience in production operations, incident response, root cause analysis, and reliability engineering.
  • Experience automating operational processes and reducing manual toil through engineering solutions.
  • Hands-on experience with Kubernetes and cloud platforms such as GCP, AWS, or Azure.
  • Experience with monitoring and observability tools including Splunk, Grafana, Prometheus, Datadog, or similar platforms.
  • Strong understanding of Linux systems, networking concepts, and distributed application architectures.
  • Experience with Infrastructure as Code and configuration management tools such as Terraform and Ansible.
  • Excellent analytical, troubleshooting, and problem-solving skills with the ability to perform effectively in mission-critical environments.

Responsibilities

  • Develop Python-based automation solutions to eliminate manual operational tasks and improve efficiency.
  • Support and maintain large-scale production systems, ensuring high availability and reliability.
  • Participate in incident response, troubleshooting, root cause analysis, and problem remediation activities.
  • Automate infrastructure management across cloud, Kubernetes, Linux, and Windows environments.
  • Implement and support infrastructure automation using tools such as Terraform and Ansible.
  • Build and maintain observability solutions including dashboards, alerts, metrics, logs, and monitoring frameworks.
  • Drive operational improvements to reduce recurring incidents and improve system stability.
  • Perform performance analysis, capacity planning, and system health assessments.
  • Support disaster recovery, failover testing, and operational readiness initiatives.
  • Evaluate and implement emerging observability, automation, and AIOps capabilities.

Job description

Job Title: Site Reliability Engineer (SRE)

Duration (Contract): 12 Months

Client Location: Southlake, TX

Location Preference: Onsite

Job Description:

As a Site Reliability Engineer (SRE), you will be responsible for improving the reliability, scalability, and operational efficiency of production systems through automation, observability, and incident management. The ideal candidate will have strong Python development expertise, hands-on experience supporting large-scale production environments, and a proven track record of reducing operational toil through automation.

Key Responsibilities:
  • Develop Python-based automation solutions to eliminate manual operational tasks and improve efficiency.
  • Support and maintain large-scale production systems, ensuring high availability and reliability.
  • Participate in incident response, troubleshooting, root cause analysis, and problem remediation activities.
  • Automate infrastructure management across cloud, Kubernetes, Linux, and Windows environments.
  • Implement and support infrastructure automation using tools such as Terraform and Ansible.
  • Build and maintain observability solutions including dashboards, alerts, metrics, logs, and monitoring frameworks.
  • Drive operational improvements to reduce recurring incidents and improve system stability.
  • Perform performance analysis, capacity planning, and system health assessments.
  • Support disaster recovery, failover testing, and operational readiness initiatives.
  • Evaluate and implement emerging observability, automation, and AIOps capabilities.
Required Skills, Experiences, Education, and Competencies:
  • Bachelor's degree in Computer Science, Engineering, or a related field, or equivalent professional experience.
  • 3-5 years of hands-on Site Reliability Engineering or Production Engineering experience supporting large-scale production systems.
  • Strong Python programming expertise with demonstrated experience building automation tools, frameworks, and operational solutions.
  • Proven experience in production operations, incident response, root cause analysis, and reliability engineering.
  • Experience automating operational processes and reducing manual toil through engineering solutions.
  • Hands-on experience with Kubernetes and cloud platforms such as GCP, AWS, or Azure.
  • Experience with monitoring and observability tools including Splunk, Grafana, Prometheus, Datadog, or similar platforms.
  • Strong understanding of Linux systems, networking concepts, and distributed application architectures.
  • Experience with Infrastructure as Code and configuration management tools such as Terraform and Ansible.
  • Excellent analytical, troubleshooting, and problem-solving skills with the ability to perform effectively in mission-critical environments.

The hourly range for roles of this nature are $40.00 to $80.00/hr. Rates are heavily dependent on skills, experience, location, and industry.

cyberThink is an Equal Opportunity Employer.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Site Reliability Engineer
Site Reliability Engineer

Spectraforce Technologies • Austin (TX)

On-site
USD 120,000 - 155,000
Site Reliability Engineer (SRE)
Site Reliability Engineer (SRE)

Knack Solutions • Reston (VA)

On-site
USD 120,000 - 160,000
SRE - Site Reliability Engineer - Senior
SRE - Site Reliability Engineer - Senior

ManpowerGroup Global, Inc. • Austin (TX)

On-site
USD 66,000 - 90,000
Site Reliability Engineer (SRE)
Site Reliability Engineer (SRE)

Piper Companies • United States

Remote
USD 120,000 - 145,000
Health insurance
Vision insurance
Dental insurance
+1
Site Reliability Engineer
Site Reliability Engineer

Ethos Group • Irving (TX)

On-site
USD 110,000 - 160,000
Site Reliability Engineer (SRE)
Site Reliability Engineer (SRE)

Randstad Digital Americas • Plano (TX)

On-site
USD 115,000 - 125,000
Medical insurance
401K plan
Sr SRE Automation Engineer
Sr SRE Automation Engineer

Compunnel, Inc. • Austin (TX), Northern (KY)

On-site
USD 130,000 - 180,000
Site Reliability Engineering (SRE)
Site Reliability Engineering (SRE)

Weekday (YC W21) • New York (NY)

On-site
USD 150,000 - 250,000
Health, dental, vision insurance
Generous PTO
Learning & development
+2
SRE Engineer
SRE Engineer

Tata Consultancy Services • Englewood Cliffs (NJ)

On-site
USD 110,000 - 125,000
Site Reliability Engineer (SRE)
Site Reliability Engineer (SRE)

myBridge Corporation • Austin (TX)

On-site
USD 120,000 - 160,000