Site Reliability Engineer (SRE) - W2 - USC OR GC
NUVENTO SYSTEMS PRIVATE LIMITED
Job description
Job Title:
Site Reliability Engineer (SRE)
Location:
Kansas City, MO (Onsite)
GC or USC ONLY
Essential Job Functions
-
Deploy, configure, and manage Azure SRE Agent across production workloads to automate incident investigation, root cause analysis, and remediation workflows
-
Monitor production systems and respond to incidents in partnership with Azure SRE Agent, reviewing agent-proposed diagnoses and mitigations and approving actions per governance policy
-
Build and maintain runbooks, subagents, and agent hooks within Azure SRE Agent to automate common operational tasks and reduce manual toil
-
Connect Azure SRE Agent to observability, incident management, and source control tooling (Azure Monitor, Application Insights, Log Analytics, GitHub, PagerDuty, or similar) to enable end-to-end automated investigations
-
Define and maintain service level objectives (SLOs), service level indicators (SLIs), and error budgets for critical services, and use agent-driven insights to track and improve them
-
Establish and enforce tool permissions, hooks, and governance controls for AI agent actions to ensure safe, auditable automation with appropriate human approval gates
-
Analyze incident trends and agent-surfaced institutional knowledge to identify recurring issues and drive permanent fixes and reliability improvements
-
Participate in on-call rotation, leveraging Azure SRE Agent to accelerate triage, reduce mean time to detect/resolve (MTTD/MTTR), and minimize after-hours disruptions
-
Author and refine incident response plans, playbooks, and postmortem processes, incorporating agent-generated documentation and learnings
-
Partner with DevOps, Platform Engineering, and development teams to instrument applications and infrastructure for effective AI-driven monitoring and diagnostics
-
Continuously evaluate new Azure SRE Agent capabilities (connectors, private plugins, subagents) and pilot adoption to expand automated coverage
-
Contribute to Infrastructure as Code and automation scripts that support reliable, repeatable, and agent-manageable environments
Experience in the following areas is required:
-
Bachelor’s degree in computer science, Information Technology, or a related field required, or equivalent professional experience
-
5+ years of experience in Site Reliability Engineering, DevOps, or a related production operations role
-
Hands-on experience with Azure SRE Agent or comparable AI-powered incident response/observability tooling strongly preferred
-
Strong working knowledge of Azure cloud services, including AKS, App Service, Azure Functions, and Azure networking
-
Experience with observability and monitoring platforms (Azure Monitor, Application Insights, Log Analytics, or similar)
-
Practical understanding of SRE fundamentals, including SLOs/SLIs, error budgets, incident management, and blameless postmortems
-
Familiarity with CI/CD tooling such as Azure DevOps Pipelines and GitHub Actions
-
Experience with Infrastructure as Code tools such as Bicep or Terraform
-
Comfort working with AI agents and automation frameworks, including reviewing and approving AI-proposed remediations under governance controls
-
Strong scripting and automation skills (PowerShell, Bash, or Python)
-
Excellent troubleshooting skills and the ability to remain calm and methodical during production incidents
-
Strong communication skills, with the ability to document findings clearly for both technical and non-technical audiences
-
Willingness to participate in an on-call rotation.