SUPERBOT (BOT) logo

Site Reliability Engineer

SUPERBOT (BOT)

On-site๐Ÿ‡ฆ๐Ÿ‡ชDubai, DU, UAEseniorPosted 13d ago

Visa & sponsorship

  • Employers in UAE sponsor the residence visa by default, and nothing in the posting says otherwise.

Job description

About The Role

Youโ€™ll sit at the intersection of software engineering and infrastructure, partnering with product and platform teams to keep our systems observable, scalable, and resilient. From shaping our on-call culture to driving infrastructure-as-code adoption, youโ€™ll have real influence over how BOT builds and operates software at scale โ€” for 500+ people and growing.

Key Responsibilities

  • Own reliability across critical services โ€” define and track SLIs/SLOs/SLAs, lead blameless post-mortems, and drive down MTTR through systematic incident management.

  • Design, build, and maintain cloud infrastructure on AWS (primary) using Terraform, ensuring environments are reproducible, version-controlled, and auditable.

  • Scale and optimize our Kubernetes-based container platform โ€” capacity planning, resource tuning, autoscaling, and cluster lifecycle management.

  • Strengthen observability end-to-end: instrument services with Prometheus and Grafana, build actionable alerting, and reduce alert noise through continuous tuning.

  • Accelerate CI/CD pipelines (GitHub Actions / GitLab CI) to support frequent, safe deployments โ€” shift reliability left by embedding checks into the delivery workflow.

  • Partner with engineering teams as an internal reliability advisor โ€” run game days, advocate for SRE best practices, and help developers build services that operate well from the start.

Job Requirements

  • 4+ years in an SRE, DevOps, or platform engineering role with production ownership at scale.

  • Hands-on Kubernetes experience โ€” deployment, scaling, networking, and troubleshooting in a production environment.

  • Infrastructure-as-code fluency with Terraform (or Pulumi) across a major cloud provider (AWS preferred).

  • Observability stack experience โ€” Prometheus, Grafana, and/or equivalent tools with a track record of building meaningful dashboards and alerts.

  • Proficiency in at least one scripting or programming language (Python, Go, or Bash) for automation and tooling.

  • AWS certification (Solutions Architect, DevOps Engineer, or SysOps Administrator).

What We Offer

  • Competitive Compensation: Enjoy a salary package tailored to your skills and experience, along with performance-based bonuses.

  • Comprehensive Benefits: We support your well-being with accommodation, meal allowances, and assistance with work visa processing.

  • Work-Life Balance: Unwind with generous holiday and New Year bonuses.

  • Top-Tier Equipment: Stay productive with the latest tools, including a MacBook and iPhone.

  • Thriving Culture: Immerse yourself in a dynamic, inclusive work environment that fosters growth.