Software Engineer SRE

The Mice Groups, Inc.

Austin (TX)

On-site

USD 83,000 - 96,000

Full time

13 days ago
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Job summary

The Mice Groups, Inc. in Austin, TX is seeking a Sr. Data Center Site Reliability Engineer to automate operations and maximize uptime of data center and facility infrastructure.

You will manage, monitor, and optimize server reliability and critical power/cooling systems. Responsibilities include building observability with Grafana, Prometheus, and Splunk; developing automation scripts in Python and Shell; maintaining NetBox DCIM data; leading incident response; and creating runbooks for standard

Qualifications

  • 8+ years of experience in site reliability engineering, production operations, or data center infrastructure management operations.
  • Bachelor's Degree in Computer Science, Computer Engineering, or a related technical field is highly preferred.
  • Deep hands-on experience troubleshooting, provisioning, and managing enterprise power, bare-metal hardware and server architectures.
  • Strong proficiency using NetBox (or similar DCIM tools) for managing rack space, device lifecycle, and asset tracking.
  • Extensive SQL experience (e.g., PostgreSQL, MySQL) for querying relational data infrastructure and deep familiarity consuming/building RESTful APIs to integrate infrastructure tools.
  • Strong expertise with Prometheus for metrics collection, Grafana for visualization, and Splunk for enterprise logging.
  • Proficient with IPMI and server out-of-band management protocols, alongside a strong understanding of data center PDU management and power feed architecture.
  • Practical understanding of data center physical infrastructure, power distribution and cooling systems.
  • Strong scripting capabilities (Python, Shell) and experience managing infrastructure across highly distributed on-premise environments.

Responsibilities

  • Enhance data center observability, logging, and alerting solutions using Grafana, Splunk, and Prometheus, building dashboards that correlate server health, network telemetry, facility power and cooling performance.
  • Develop automation scripts for hardware incident triage, alert noise reduction, log correlation, and operational workflows, converting recurring manual bare-power/cooling infrastructure investigation patterns into reusable tooling.
  • Maintain our NetBox data center inventory, building automated pipelines via APIs to track physical infrastructure, rack layouts, and cable topologies.
  • Build and tune Grafana dashboards with complex queries spanning multiple data sources (including Prometheus metrics) for server health visualization, bare-metal hardware bottleneck identification, and data center capacity monitoring using power feed and cooling infrastructure metrics.
  • Utilize Splunk and relational databases for infrastructure analytics, writing extensive SQL queries and SPL queries to troubleshoot server production issues, identify infrastructure bottlenecks, and surface environmental insights via IPMI interfaces into dashboards.
  • Lead incident response and on-call rotations for high-severity data center infrastructure events, directing triage, root cause analysis, mitigation, and resolution for both server-level and facility-level power feed or environmental anomalies.
  • Develop and maintain runbooks, hardware operational playbooks, and process documentation for common facility, power feed, and server failure scenarios, standardizing infrastructure SOPs across the SRE organization.
  • Collaborate closely with development, hardware engineering, and facility operations teams to integrate observability best practices into the infrastructure lifecycle and embed monitoring into new compute, storage, power and cooling system rollouts.

Skills

SRE experience
Incident response
Automation scripting
Monitoring & observability
Data center operations

Education

Bachelor's Degree in Computer Science/Engineering

Tools

Grafana
Prometheus
Splunk
NetBox
PostgreSQL
MySQL
RESTful APIs
IPMI
Python
Shell

Job description

Location: Austin, TX (On Site 40 hours per week)

Contract Duration: 12+ month

Pay Rate: $60-$70/hourly (W2)

Job Description

We are seeking an experienced Sr. Data Center Site Reliability Engineer to automate operations and maximize the uptime, efficiency, and scalability of data center, facility power/cooling infrastructure, and software automation. In this role, you will manage, monitor, and optimizing both server reliability and the critical power and cooling infrastructure that sustains our distributed production systems.

Key Responsibilities
  • Enhance data center observability, logging, and alerting solutions using Grafana, Splunk, and Prometheus, building dashboards that correlate server health, network telemetry, facility power and cooling performance.
  • Develop automation scripts for hardware incident triage, alert noise reduction, log correlation, and operational workflows, converting recurring manual bare-power/cooling infrastructure investigation patterns into reusable tooling.
  • Maintain our NetBox data center inventory, building automated pipelines via APIs to track physical infrastructure, rack layouts, and cable topologies.
  • Build and tune Grafana dashboards with complex queries spanning multiple data sources (including Prometheus metrics) for server health visualization, bare-metal hardware bottleneck identification, and data center capacity monitoring using power feed and cooling infrastructure metrics.
  • Utilize Splunk and relational databases for infrastructure analytics, writing extensive SQL queries and SPL queries to troubleshoot server production issues, identify infrastructure bottlenecks, and surface environmental insights via IPMI interfaces into dashboards.
  • Lead incident response and on-call rotations for high-severity data center infrastructure events, directing triage, root cause analysis, mitigation, and resolution for both server-level and facility-level power feed or environmental anomalies.
  • Develop and maintain runbooks, hardware operational playbooks, and process documentation for common facility, power feed, and server failure scenarios, standardizing infrastructure SOPs across the SRE organization.
  • Collaborate closely with development, hardware engineering, and facility operations teams to integrate observability best practices into the infrastructure lifecycle and embed monitoring into new compute, storage, power and cooling system rollouts.
Qualifications
  • Experience: 8+ years of experience in site reliability engineering, production operations, or data center infrastructure management operations.
  • Education: Bachelor's Degree in Computer Science, Computer Engineering, or a related technical field is highly preferred.
  • Bare-Metal & Hardware Expertise: Deep hands-on experience troubleshooting, provisioning, and managing enterprise power, bare-metal hardware and server architectures.
  • Inventory & Asset Management: Strong proficiency using NetBox (or similar DCIM tools) for managing rack space, device lifecycle, and asset tracking.
  • Data & API Capabilities: Extensive SQL experience (e.g., PostgreSQL, MySQL) for querying relational data infrastructure and deep familiarity consuming/building RESTful APIs to integrate infrastructure tools.
  • Monitoring & Tooling: Strong expertise with Prometheus for metrics collection, Grafana for visualization, and Splunk for enterprise logging.
  • Infrastructure Protocols: Proficient with IPMI and server out-of-band management protocols, alongside a strong understanding of data center PDU management and power feed architecture.
  • Facilities Knowledge: Practical understanding of data center physical infrastructure, specifically power feed distribution systems and cooling infrastructure (e.g., HVAC, liquid cooling, hot/cold aisle containment, air handling units).
  • Automation: Strong scripting capabilities (Python, Shell) and experience managing infrastructure across highly distributed on-premise environments.

We are an equal opportunity employer and value diversity at The Mice Groups Inc. We do not discriminate on the basis of race, religion, color, national origin, gender, sexual orientation, age, marital status, veteran status, or disability status.

Pursuant to the San Francisco Fair Chance Ordinance, we will consider for employment qualified applicants with arrest and conviction records.

Pursuant to the Los Angeles Fair Chance Ordinance, we will consider for employment qualified applicants with arrest and conviction records.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Site Reliability Engineer
Site Reliability Engineer

The Mice Groups, Inc. • Austin (TX)

On-site
USD 140,000 - 190,000
Senior Data Center SRE - Automation & Uptime Lead
Senior Data Center SRE - Automation & Uptime Lead

The Mice Groups, Inc. • Austin (TX)

On-site
USD 83,000 - 96,000
Site Reliability Engineer (SRE)
Site Reliability Engineer (SRE)

Smart IMS Inc • Southlake (TX)

On-site
USD 55,000 - 110,000
Site Reliability Engineer
Site Reliability Engineer

Bolt Graphics, Inc. • Sunnyvale (CA)

On-site
USD 145,000 - 165,000
100% covered medical, dental, and vision premiums
Equity - Stock Options
401(k) match
Site Reliability Engineer
Site Reliability Engineer

ltdglobal • Berkeley (CA)

Hybrid
USD 96,000 - 124,000
SRE - Site Reliability Engineer - Senior
SRE - Site Reliability Engineer - Senior

ManpowerGroup Global, Inc. • Austin (TX)

On-site
USD 66,000 - 90,000
Data Center Engineer
Data Center Engineer

Motion Recruitment • Oklahoma

On-site
USD 80,000 - 120,000
Medical Insurance
401(k) including match
Health Savings Account
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Spectraforce Technologies • Austin (TX)

On-site
USD 130,000 - 170,000
Sr Data Center Engineer
Sr Data Center Engineer

Randstad Digital Americas • Newport Beach (CA)

On-site
USD 110,000 - 135,000
Site Reliability Engineer
Site Reliability Engineer

Ltd Global • Berkeley (CA)

On-site
USD 94,000 - 127,000