Evlo AI logo

DevOps Engineer

Evlo AI

On-siteSan Diego, CAmidPosted 14h ago

Job description

About The Role

The DevOps Engineer will build and operate the infrastructure, deployment systems, and observability platforms that keep production services reliable at scale. The role spans AWS cloud infrastructure, Kubernetes, CI/CD, infrastructure as code, and incident response across distributed systems.

The engineer will partner with software development and security teams to improve release velocity without compromising availability or operational safety. Success means repeatable deployments, clear service ownership, measurable reliability, and infrastructure that scales with product demand.

Key Responsibilities

  • Design and maintain highly available AWS infrastructure using Terraform, including VPCs, IAM, ECS or EKS, RDS, S3, and CloudFront

  • Build and improve CI/CD pipelines with GitHub Actions, GitLab CI, or Jenkins, including automated testing, security scanning, progressive delivery, and rollback procedures

  • Operate Kubernetes-based workloads, managing deployments, ingress, autoscaling, secrets, resource limits, and cluster upgrades

  • Implement observability using Prometheus, Grafana, CloudWatch, OpenTelemetry, and centralized logging to improve detection and resolution of production issues

  • Automate operational workflows with Python, Go, or Bash, reducing manual intervention in provisioning, deployments, incident response, and routine maintenance

  • Define and track reliability practices including SLIs, SLOs, error budgets, capacity planning, and disaster recovery testing

  • Participate in on-call rotations, lead incident response, document root-cause analyses, and deliver follow-up improvements that prevent recurrence

What We Are Looking For

  • 3–8 years of experience in DevOps, site reliability engineering, platform engineering, or a closely related infrastructure role

  • Hands-on experience operating production workloads in AWS, with strong knowledge of networking, IAM, compute, storage, databases, and security controls

  • Proficiency with Terraform or an equivalent infrastructure-as-code tool and practical experience managing infrastructure through version-controlled workflows

  • Experience with Docker and Kubernetes, including workload scheduling, service discovery, ingress, autoscaling, and troubleshooting

  • Strong understanding of CI/CD design, Git-based development workflows, automated testing, release strategies, and deployment rollback patterns

  • Bachelor’s degree in computer science, engineering, information systems, or equivalent practical experience

  • Bonus: Experience with Go, Argo CD, Helm, service mesh technologies, OpenTelemetry, PostgreSQL operations, compliance automation, or multi-region disaster recovery