Evlo AI logo

Site Reliability Engineer

Evlo AI

RemotemidPosted 2h ago

Job description

About The Role

The Site Reliability Engineer will design, automate, and operate the infrastructure supporting production services at scale. The role focuses on Kubernetes-based platforms, cloud networking, observability, incident response, and the reliability of distributed systems running across AWS environments.

Working with application engineers, security, and platform teams, the engineer will turn operational requirements into resilient systems and repeatable delivery workflows. The role is remote from New York, NY, with direct ownership of service-level objectives, production readiness, and continuous improvements to availability and performance.

Key Responsibilities

  • Design and operate highly available production infrastructure across AWS using Kubernetes, Terraform, and infrastructure-as-code best practices

  • Build and maintain CI/CD pipelines with tools such as GitHub Actions, Argo CD, or Jenkins to automate testing, deployments, rollbacks, and environment provisioning

  • Define and enforce service-level objectives, error budgets, and production readiness standards for critical services

  • Implement observability using Prometheus, Grafana, OpenTelemetry, and centralized logging platforms to reduce detection and resolution time

  • Lead incident response, coordinate technical remediation, and produce blameless postmortems with measurable follow-up actions

  • Improve platform resilience through capacity planning, performance testing, disaster recovery exercises, and automated failure remediation

  • Partner with development teams to improve application architecture, containerization, security controls, and operational runbooks

What We Are Looking For

  • 3โ€“8 years of experience in site reliability engineering, DevOps, platform engineering, or a related production infrastructure role

  • Strong hands-on experience with AWS services such as EC2, EKS, IAM, VPC, S3, RDS, and CloudWatch

  • Proficiency with Kubernetes, Docker, Terraform, and Linux systems administration in production environments

  • Experience building and operating CI/CD pipelines and managing deployment strategies such as canary releases, blue-green deployments, and automated rollbacks

  • Working knowledge of observability, distributed systems, networking, and incident management practices, including SLOs, SLIs, and error budgets

  • Proficiency in Python, Go, or Bash for automation, tooling, and operational workflows; bachelorโ€™s degree in computer science, engineering, or equivalent practical experience

  • Bonus: Experience with Argo CD, Helm, Prometheus, Grafana, OpenTelemetry, service meshes, multi-region architectures, or chaos engineering