N

Site Reliability Engineer (SRE) - W2 - USC OR GC

NUVENTO SYSTEMS PRIVATE LIMITED

On-siteKansas City, MOseniorPosted 1h ago

Job description

Job Title:

Site Reliability Engineer (SRE)

Location:

Kansas City, MO (Onsite)

GC or USC ONLY

Essential Job Functions

  • Deploy, configure, and manage Azure SRE Agent across production workloads to automate incident investigation, root cause analysis, and remediation workflows

  • Monitor production systems and respond to incidents in partnership with Azure SRE Agent, reviewing agent-proposed diagnoses and mitigations and approving actions per governance policy

  • Build and maintain runbooks, subagents, and agent hooks within Azure SRE Agent to automate common operational tasks and reduce manual toil

  • Connect Azure SRE Agent to observability, incident management, and source control tooling (Azure Monitor, Application Insights, Log Analytics, GitHub, PagerDuty, or similar) to enable end-to-end automated investigations

  • Define and maintain service level objectives (SLOs), service level indicators (SLIs), and error budgets for critical services, and use agent-driven insights to track and improve them

  • Establish and enforce tool permissions, hooks, and governance controls for AI agent actions to ensure safe, auditable automation with appropriate human approval gates

  • Analyze incident trends and agent-surfaced institutional knowledge to identify recurring issues and drive permanent fixes and reliability improvements

  • Participate in on-call rotation, leveraging Azure SRE Agent to accelerate triage, reduce mean time to detect/resolve (MTTD/MTTR), and minimize after-hours disruptions

  • Author and refine incident response plans, playbooks, and postmortem processes, incorporating agent-generated documentation and learnings

  • Partner with DevOps, Platform Engineering, and development teams to instrument applications and infrastructure for effective AI-driven monitoring and diagnostics

  • Continuously evaluate new Azure SRE Agent capabilities (connectors, private plugins, subagents) and pilot adoption to expand automated coverage

  • Contribute to Infrastructure as Code and automation scripts that support reliable, repeatable, and agent-manageable environments

Experience in the following areas is required:

  • Bachelor’s degree in computer science, Information Technology, or a related field required, or equivalent professional experience

  • 5+ years of experience in Site Reliability Engineering, DevOps, or a related production operations role

  • Hands-on experience with Azure SRE Agent or comparable AI-powered incident response/observability tooling strongly preferred

  • Strong working knowledge of Azure cloud services, including AKS, App Service, Azure Functions, and Azure networking

  • Experience with observability and monitoring platforms (Azure Monitor, Application Insights, Log Analytics, or similar)

  • Practical understanding of SRE fundamentals, including SLOs/SLIs, error budgets, incident management, and blameless postmortems

  • Familiarity with CI/CD tooling such as Azure DevOps Pipelines and GitHub Actions

  • Experience with Infrastructure as Code tools such as Bicep or Terraform

  • Comfort working with AI agents and automation frameworks, including reviewing and approving AI-proposed remediations under governance controls

  • Strong scripting and automation skills (PowerShell, Bash, or Python)

  • Excellent troubleshooting skills and the ability to remain calm and methodical during production incidents

  • Strong communication skills, with the ability to document findings clearly for both technical and non-technical audiences

  • Willingness to participate in an on-call rotation.