International Blueberry Organization logo

Site Reliability Engineering Manager

International Blueberry Organization

On-site๐Ÿ‡ฆ๐Ÿ‡ชUAEleadPosted 6d ago

Visa & sponsorship

  • Employers in UAE sponsor the residence visa by default, and nothing in the posting says otherwise.

Job description

โš™๏ธ Site Reliability Engineering Manager๐Ÿš€ Role Description

We are looking for a strategic, technically strong, and people-focused

Site Reliability Engineering Manager

to lead initiatives that improve the reliability, scalability, availability, and performance of critical software systems. ๐ŸŒ In this role, you will guide reliability engineering practices, promote automation, and help build resilient infrastructure and services that deliver a consistent user experience.

You will work closely with software engineering, DevOps, infrastructure, security, and product teams to establish reliability goals, monitor system health, and identify opportunities for improving operational performance. ๐Ÿ“Š You will support incident management, root cause analysis, capacity planning, disaster recovery, and service-level objectives (SLOs).

You will also drive automation and operational efficiency by improving deployment processes, observability, monitoring, alerting, and infrastructure management. ๐Ÿ› ๏ธ A key part of the role will be creating a culture of reliability, continuous improvement, knowledge sharing, and effective technical collaboration.

The ideal candidate should combine strong technical judgment with excellent leadership, communication, and problem-solving abilities. ๐Ÿง  You should be comfortable making data-driven decisions, managing operational priorities, and guiding teams through complex reliability challenges.

๐ŸŽฏ Qualifications

  • ๐ŸŽ“ Bachelorโ€™s degree or equivalent qualification in Computer Science, Software Engineering, Information Technology, or a related field.

  • โš™๏ธ Strong understanding of Site Reliability Engineering (SRE), DevOps, cloud infrastructure, and distributed systems.

  • โ˜๏ธ Familiarity with cloud platforms such as AWS, Azure, or Google Cloud.

  • ๐Ÿณ Knowledge of containers, Kubernetes, CI/CD pipelines, infrastructure as code, and automation tools.

  • ๐Ÿ“Š Strong understanding of monitoring, observability, logging, alerting, system performance, and service reliability metrics.

  • ๐Ÿšจ Knowledge of incident management, root cause analysis, disaster recovery, and capacity planning.

  • ๐Ÿ” Understanding of security, scalability, availability, and resilient system design principles.

  • ๐Ÿง  Strong technical problem-solving and analytical skills with a focus on continuous improvement.

  • ๐Ÿค Excellent leadership, communication, mentoring, and cross-functional collaboration skills.

  • ๐Ÿ“ˆ Ability to define and monitor SLOs, SLIs, operational metrics, and reliability objectives.

  • ๐Ÿ“ Strong documentation and decision-making skills with a structured approach to complex technical challenges.

  • ๐Ÿ’ก Proactive mindset with a passion for automation, operational excellence, system reliability, and engineering innovation.