Job Description
We are seeking an experienced Site Reliability Engineer (SRE) with strong data center and bare-metal infrastructure experience. This role focuses on maintaining and improving the reliability of data center infrastructure through monitoring, automation, troubleshooting, and incident response.
You will work across servers, hardware, power/cooling infrastructure, monitoring, and automation, partnering with infrastructure, hardware, and data center operations teams to improve reliability and operational efficiency.
Key Responsibilities
- Monitor and improve data center infrastructure using Prometheus, Grafana, and Splunk, including server health and power/cooling telemetry.
- Develop Python and Shell scripts to automate troubleshooting, incident response, alert management, and other operational processes.
- Maintain and improve NetBox or similar data center inventory systems, including device, rack, and infrastructure information.
- Build and maintain Grafana dashboards to monitor server health, infrastructure performance, capacity, and power/cooling metrics.
- Use SQL and Splunk queries to troubleshoot infrastructure issues, analyze system data, and identify potential bottlenecks.
- Troubleshoot bare-metal servers, hardware, and data center infrastructure, including issues involving IPMI, PDUs, and power feeds.
- Participate in incident response and on-call support, including troubleshooting, root cause analysis, mitigation, and resolution of infrastructure issues.
- Develop and maintain runbooks and operational documentation for common server, hardware, power, cooling, and facility-related issues.
- Work with software, hardware, and data center operations teams to support reliable infrastructure deployments and ongoing operations.
Qualifications
- 8+ years of experience in SRE, Production Operations, Data Center Infrastructure, or a similar role.
- Strong hands‑on experience with bare‑metal servers, hardware troubleshooting, provisioning, and data center infrastructure.
- Experience with NetBox or similar DCIM/inventory tools.
- Strong SQL experience and experience working with REST APIs.
- Hands‑on experience with Prometheus, Grafana, and Splunk.
- Strong understanding of IPMI, out‑of‑band server management, PDUs, and power distribution.
- Practical understanding of data center power and cooling systems, including HVAC, liquid cooling, hot/cold aisle containment, and air handling.
- Strong Python and Shell scripting skills with experience automating on‑premises infrastructure.
- Bachelor’s degree in Computer Science, Computer Engineering, or a related technical field preferred.