
Site Reliability Engineer
The Mice Groups
Job description
Job Description
We are seeking an experienced Site Reliability Engineer (SRE) with strong data center and bare-metal infrastructure experience. This role focuses on maintaining and improving the reliability of data center infrastructure through monitoring, automation, troubleshooting, and incident response.
You will work across servers, hardware, power/cooling infrastructure, monitoring, and automation
, partnering with infrastructure, hardware, and data center operations teams to improve reliability and operational efficiency.
Key Responsibilities
-
Monitor and improve data center infrastructure using Prometheus, Grafana, and Splunk
, including server health and power/cooling telemetry.
-
Develop Python and Shell scripts to automate troubleshooting, incident response, alert management, and other operational processes.
-
Maintain and improve NetBox or similar data center inventory systems, including device, rack, and infrastructure information.
-
Build and maintain Grafana dashboards to monitor server health, infrastructure performance, capacity, and power/cooling metrics.
-
Use SQL and Splunk queries to troubleshoot infrastructure issues, analyze system data, and identify potential bottlenecks.
-
Troubleshoot bare-metal servers, hardware, and data center infrastructure
, including issues involving IPMI, PDUs, and power feeds.
-
Participate in incident response and on-call support
, including troubleshooting, root cause analysis, mitigation, and resolution of infrastructure issues.
-
Develop and maintain runbooks and operational documentation for common server, hardware, power, cooling, and facility-related issues.
-
Work with software, hardware, and data center operations teams to support reliable infrastructure deployments and ongoing operations.
Qualifications
-
8+ years of experience in SRE, Production Operations, Data Center Infrastructure, or a similar role.
-
Strong hands-on experience with bare-metal servers, hardware troubleshooting, provisioning, and data center infrastructure
.
-
Experience with NetBox or similar DCIM/inventory tools
.
-
Strong SQL experience and experience working with REST APIs
.
-
Hands-on experience with Prometheus, Grafana, and Splunk
.
-
Strong understanding of IPMI, out-of-band server management, PDUs, and power distribution
.
-
Practical understanding of data center power and cooling systems
, including HVAC, liquid cooling, hot/cold aisle containment, and air handling.
-
Strong Python and Shell scripting skills with experience automating on-premises infrastructure.
-
Bachelor's degree in Computer Science, Computer Engineering, or a related technical field preferred
.