Senior Data Center SRE - Automation & Uptime Lead

The Mice Groups, Inc.

Austin (TX)

On-site

USD 83,000 - 96,000

Full time

13 days ago
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Job summary

The Mice Groups, Inc. in Austin, TX is seeking a Sr. Data Center Site Reliability Engineer to automate operations and maximize uptime of data center and facility infrastructure.

You will manage, monitor, and optimize server reliability and critical power/cooling systems. Responsibilities include building observability with Grafana, Prometheus, and Splunk; developing automation scripts in Python and Shell; maintaining NetBox DCIM data; leading incident response; and creating runbooks for standard

Qualifications

  • 8+ years of experience in site reliability engineering, production operations, or data center infrastructure management operations.
  • Bachelor's Degree in Computer Science, Computer Engineering, or a related technical field is highly preferred.
  • Deep hands-on experience troubleshooting, provisioning, and managing enterprise power, bare-metal hardware and server architectures.
  • Strong proficiency using NetBox (or similar DCIM tools) for managing rack space, device lifecycle, and asset tracking.
  • Extensive SQL experience (e.g., PostgreSQL, MySQL) for querying relational data infrastructure and deep familiarity consuming/building RESTful APIs to integrate infrastructure tools.
  • Strong expertise with Prometheus for metrics collection, Grafana for visualization, and Splunk for enterprise logging.
  • Proficient with IPMI and server out-of-band management protocols, alongside a strong understanding of data center PDU management and power feed architecture.
  • Practical understanding of data center physical infrastructure, power distribution and cooling systems.
  • Strong scripting capabilities (Python, Shell) and experience managing infrastructure across highly distributed on-premise environments.

Responsibilities

  • Enhance data center observability, logging, and alerting solutions using Grafana, Splunk, and Prometheus, building dashboards that correlate server health, network telemetry, facility power and cooling performance.
  • Develop automation scripts for hardware incident triage, alert noise reduction, log correlation, and operational workflows, converting recurring manual bare-power/cooling infrastructure investigation patterns into reusable tooling.
  • Maintain our NetBox data center inventory, building automated pipelines via APIs to track physical infrastructure, rack layouts, and cable topologies.
  • Build and tune Grafana dashboards with complex queries spanning multiple data sources (including Prometheus metrics) for server health visualization, bare-metal hardware bottleneck identification, and data center capacity monitoring using power feed and cooling infrastructure metrics.
  • Utilize Splunk and relational databases for infrastructure analytics, writing extensive SQL queries and SPL queries to troubleshoot server production issues, identify infrastructure bottlenecks, and surface environmental insights via IPMI interfaces into dashboards.
  • Lead incident response and on-call rotations for high-severity data center infrastructure events, directing triage, root cause analysis, mitigation, and resolution for both server-level and facility-level power feed or environmental anomalies.
  • Develop and maintain runbooks, hardware operational playbooks, and process documentation for common facility, power feed, and server failure scenarios, standardizing infrastructure SOPs across the SRE organization.
  • Collaborate closely with development, hardware engineering, and facility operations teams to integrate observability best practices into the infrastructure lifecycle and embed monitoring into new compute, storage, power and cooling system rollouts.

Skills

SRE experience
Incident response
Automation scripting
Monitoring & observability
Data center operations

Education

Bachelor's Degree in Computer Science/Engineering

Tools

Grafana
Prometheus
Splunk
NetBox
PostgreSQL
MySQL
RESTful APIs
IPMI
Python
Shell

Job description

The Mice Groups, Inc. in Austin, TX is seeking a Sr. Data Center Site Reliability Engineer to automate operations and maximize uptime of data center and facility infrastructure.

You will manage, monitor, and optimize server reliability and critical power/cooling systems. Responsibilities include building observability with Grafana, Prometheus, and Splunk; developing automation scripts in Python and Shell; maintaining NetBox DCIM data; leading incident response; and creating runbooks for standard

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Data Center SRE - Bare-Metal Reliability & Automation
Data Center SRE - Bare-Metal Reliability & Automation

The Mice Groups, Inc. • Austin (TX)

On-site
USD 140,000 - 190,000
Software Engineer SRE
Software Engineer SRE

The Mice Groups, Inc. • Austin (TX)

On-site
USD 83,000 - 96,000
Site Reliability Engineer
Site Reliability Engineer

The Mice Groups, Inc. • Austin (TX)

On-site
USD 140,000 - 190,000
SRE - Site Reliability Engineer - Senior
SRE - Site Reliability Engineer - Senior

ManpowerGroup Global, Inc. • Austin (TX)

On-site
USD 66,000 - 90,000
Sr. Site Reliability Engineer - SRE, Onsite - 70170
Sr. Site Reliability Engineer - SRE, Onsite - 70170

PRIMUS Global Services • Town of Texas (WI)

On-site
USD 140,000 - 190,000
24/7 Site Reliability Engineer for HPC Data Center
24/7 Site Reliability Engineer for HPC Data Center

BCP Engineers & Consultants • Berkeley (CA)

On-site
USD 110,000 - 170,000
Senior SRE — Scale, Automation & Uptime
Senior SRE — Scale, Automation & Uptime

Hirebridge • Northern (KY)

Hybrid
USD 110,000 - 145,000
Bonus
Senior Lead SRE: Drive Automation & Reliability (Remote)
Senior Lead SRE: Drive Automation & Reliability (Remote)

Akamai Career Site • United States

Remote
USD 121,000 - 219,000
Healthcare benefits
Equity awards
Employee Stock Purchase Plan (ESPP)
Senior Data Center Infrastructure Engineer - High-Uptime
Senior Data Center Infrastructure Engineer - High-Uptime

Setec Alpha • New York (NY)

On-site
USD 120,000 - 160,000
Data Center Building Engineer Lead: Reliability & Operations
Data Center Building Engineer Lead: Reliability & Operations

NextGenEnergyJobs • Austin (TX)

On-site
USD 140,000 - 165,000
Health benefits
401(k) match
Paid time off
+1