LTD Global logo

Site Reliability Engineer

LTD Global

HybridBerkeley, CAmid$156k–$166kPosted 3h ago

Job description

Hybrid — Berkeley, CA

1 Year Contract Assignment w****ith possibility of extension based on performance and organizational needs.

$80/hr

Ever wondered what powers breakthrough research in energy, physics, materials science, and chemistry? You're looking at it. This national HPC facility supports 11,000+ scientists pushing the boundaries of what's possible, and we need a sharp, self-motivated SRE to help keep that engine running without interruption.

If you love solving real problems on live infrastructure, thrive on ownership, and want your work to directly enable world-class science, this is your seat.

What You'll Own

  • Monitor and triage alerts across compute, storage, network, and facility systems in real time

  • Build automation that prevents issues before they become outages

Develop new tools and integrations across the monitoring pipeline (APIs alerts- action)

  • Walk the data center floor to keep power, cooling, and environmental systems humming

  • Coordinate maintenance activities across teams and keep incidents accurately tracked

  • Dig into complex, ambiguous problems and drive them to resolution

What You Bring

  • Comfort working Owl shift (12am–8am), 5 days/week, hybrid onsite in Berkeley, CA

  • Solid Linux/command-line (SSH) chops

  • Programming/scripting experience: Python, C, C++, Perl, or Java

  • A self-starter mindset, eager to pick up Kubernetes, Prometheus/VictoriaMetrics, Alertmanager, and building management/cooling systems

  • Network security fundamentals (ACLs, firewalls)

  • Strong cross-team communication and collaboration skills

Nice to Have

  • Experience building or deploying Agentic AI / autonomous automation for technical workflows

  • ServiceNow implementation experience

  • ITSM best-practice know-how

lmuc9HhFwu