Site Reliability Engineer

The Mice Groups, Inc.

Austin (TX)

On-site

USD 140,000 - 190,000

Full time

8 days ago
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Job summary

The Mice Groups, Inc. is seeking an experienced Site Reliability Engineer (SRE) to maintain and improve the reliability of data center infrastructure, including bare‑metal servers, power and cooling, and monitoring systems.

You will automate troubleshooting with Python and Shell scripts, build Grafana dashboards, manage NetBox inventory, and use Prometheus, Grafana, and Splunk for observability and incident response.

Qualifications

  • 8+ years of experience in SRE/Production Operations/Data Center Infrastructure.
  • Hands-on with bare‑metal servers and data center infrastructure.
  • Experience with NetBox or similar DCIM/inventory tools.
  • Strong SQL experience and REST APIs.
  • Hands-on experience with Prometheus, Grafana, and Splunk.
  • Knowledge of IPMI, out‑of‑band server management, PDUs.
  • Data center power and cooling systems; HVAC, liquid cooling, containment.

Responsibilities

  • Monitor and improve data center infrastructure using Prometheus, Grafana, and Splunk.
  • Develop Python and Shell scripts to automate troubleshooting, incident response, alert management, and other operational processes.
  • Maintain and improve NetBox DCIM inventory systems, including device, rack, and infrastructure information.
  • Build and maintain Grafana dashboards to monitor server health, infrastructure performance, capacity, and power/cooling metrics.
  • Use SQL and Splunk queries to troubleshoot infrastructure issues, analyze system data, and identify bottlenecks.
  • Troubleshoot bare-metal servers, hardware, and data center infrastructure, including IPMI, PDUs, and power feeds.
  • Participate in incident response and on‑call support, including root cause analysis, mitigation, and resolution.
  • Develop and maintain runbooks and operational documentation for common server, hardware, power, cooling, and facility issues.
  • Work with software, hardware, and data center operations teams to support reliable deployments and operations.

Skills

SRE
Production Operations
Data Center Infrastructure
Python scripting
Shell scripting
SQL
REST APIs
On‑call
Incident response

Education

Bachelor’s degree preferred

Tools

NetBox
Prometheus
Grafana
Splunk
SQL
IPMI
PDUs

Job description

Job Description

We are seeking an experienced Site Reliability Engineer (SRE) with strong data center and bare-metal infrastructure experience. This role focuses on maintaining and improving the reliability of data center infrastructure through monitoring, automation, troubleshooting, and incident response.

You will work across servers, hardware, power/cooling infrastructure, monitoring, and automation, partnering with infrastructure, hardware, and data center operations teams to improve reliability and operational efficiency.

Key Responsibilities

  • Monitor and improve data center infrastructure using Prometheus, Grafana, and Splunk, including server health and power/cooling telemetry.
  • Develop Python and Shell scripts to automate troubleshooting, incident response, alert management, and other operational processes.
  • Maintain and improve NetBox or similar data center inventory systems, including device, rack, and infrastructure information.
  • Build and maintain Grafana dashboards to monitor server health, infrastructure performance, capacity, and power/cooling metrics.
  • Use SQL and Splunk queries to troubleshoot infrastructure issues, analyze system data, and identify potential bottlenecks.
  • Troubleshoot bare-metal servers, hardware, and data center infrastructure, including issues involving IPMI, PDUs, and power feeds.
  • Participate in incident response and on-call support, including troubleshooting, root cause analysis, mitigation, and resolution of infrastructure issues.
  • Develop and maintain runbooks and operational documentation for common server, hardware, power, cooling, and facility-related issues.
  • Work with software, hardware, and data center operations teams to support reliable infrastructure deployments and ongoing operations.

Qualifications

  • 8+ years of experience in SRE, Production Operations, Data Center Infrastructure, or a similar role.
  • Strong hands‑on experience with bare‑metal servers, hardware troubleshooting, provisioning, and data center infrastructure.
  • Experience with NetBox or similar DCIM/inventory tools.
  • Strong SQL experience and experience working with REST APIs.
  • Hands‑on experience with Prometheus, Grafana, and Splunk.
  • Strong understanding of IPMI, out‑of‑band server management, PDUs, and power distribution.
  • Practical understanding of data center power and cooling systems, including HVAC, liquid cooling, hot/cold aisle containment, and air handling.
  • Strong Python and Shell scripting skills with experience automating on‑premises infrastructure.
  • Bachelor’s degree in Computer Science, Computer Engineering, or a related technical field preferred.
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Software Engineer SRE
Software Engineer SRE

The Mice Groups, Inc. • Austin (TX)

On-site
USD 83,000 - 96,000
Data Center SRE - Bare-Metal Reliability & Automation
Data Center SRE - Bare-Metal Reliability & Automation

The Mice Groups, Inc. • Austin (TX)

On-site
USD 140,000 - 190,000
Site Reliability Engineer
Site Reliability Engineer

Ethos Group • Irving (TX)

On-site
USD 110,000 - 160,000
Site Reliability Engineer
Site Reliability Engineer

CT19 • Massachusetts

On-site
USD 140,000 - 210,000
Senior SRE Engineer
Senior SRE Engineer

Compunnel, Inc. • Alpharetta (GA)

On-site
USD 140,000 - 190,000
Site Reliability Engineer (SRE)
Site Reliability Engineer (SRE)

Methodic • San Francisco (CA)

On-site
USD 140,000 - 210,000
Site Reliability Engineer (SRE)
Site Reliability Engineer (SRE)

Knack Solutions • Reston (VA)

On-site
USD 120,000 - 160,000
Site Reliability Engineer (SRE)
Site Reliability Engineer (SRE)

Smart IMS Inc • Southlake (TX)

On-site
USD 55,000 - 110,000
Site Reliability Engineer
Site Reliability Engineer

TalentDome Staffing • United States

On-site
USD 140,000 - 210,000
Site Reliability Engineer
Site Reliability Engineer

myBridge Corporation • Washington, Northern (KY)

Hybrid
USD 110,000 - 140,000