The Mice Groups logo

Site Reliability Engineer

The Mice Groups

On-siteAustin, TXseniorPosted 11h ago

Job description

Job Description

We are seeking an experienced Site Reliability Engineer (SRE) with strong data center and bare-metal infrastructure experience. This role focuses on maintaining and improving the reliability of data center infrastructure through monitoring, automation, troubleshooting, and incident response.

You will work across servers, hardware, power/cooling infrastructure, monitoring, and automation

, partnering with infrastructure, hardware, and data center operations teams to improve reliability and operational efficiency.

Key Responsibilities

  • Monitor and improve data center infrastructure using Prometheus, Grafana, and Splunk

    , including server health and power/cooling telemetry.

  • Develop Python and Shell scripts to automate troubleshooting, incident response, alert management, and other operational processes.

  • Maintain and improve NetBox or similar data center inventory systems, including device, rack, and infrastructure information.

  • Build and maintain Grafana dashboards to monitor server health, infrastructure performance, capacity, and power/cooling metrics.

  • Use SQL and Splunk queries to troubleshoot infrastructure issues, analyze system data, and identify potential bottlenecks.

  • Troubleshoot bare-metal servers, hardware, and data center infrastructure

    , including issues involving IPMI, PDUs, and power feeds.

  • Participate in incident response and on-call support

    , including troubleshooting, root cause analysis, mitigation, and resolution of infrastructure issues.

  • Develop and maintain runbooks and operational documentation for common server, hardware, power, cooling, and facility-related issues.

  • Work with software, hardware, and data center operations teams to support reliable infrastructure deployments and ongoing operations.

Qualifications

  • 8+ years of experience in SRE, Production Operations, Data Center Infrastructure, or a similar role.

  • Strong hands-on experience with bare-metal servers, hardware troubleshooting, provisioning, and data center infrastructure

    .

  • Experience with NetBox or similar DCIM/inventory tools

    .

  • Strong SQL experience and experience working with REST APIs

    .

  • Hands-on experience with Prometheus, Grafana, and Splunk

    .

  • Strong understanding of IPMI, out-of-band server management, PDUs, and power distribution

    .

  • Practical understanding of data center power and cooling systems

    , including HVAC, liquid cooling, hot/cold aisle containment, and air handling.

  • Strong Python and Shell scripting skills with experience automating on-premises infrastructure.

  • Bachelor's degree in Computer Science, Computer Engineering, or a related technical field preferred

    .