
Site Reliability Engineering Manager
International Blueberry Organization
Visa & sponsorship
- Employers in UAE sponsor the residence visa by default, and nothing in the posting says otherwise.
Job description
โ๏ธ Site Reliability Engineering Manager๐ Role Description
We are looking for a strategic, technically strong, and people-focused
Site Reliability Engineering Manager
to lead initiatives that improve the reliability, scalability, availability, and performance of critical software systems. ๐ In this role, you will guide reliability engineering practices, promote automation, and help build resilient infrastructure and services that deliver a consistent user experience.
You will work closely with software engineering, DevOps, infrastructure, security, and product teams to establish reliability goals, monitor system health, and identify opportunities for improving operational performance. ๐ You will support incident management, root cause analysis, capacity planning, disaster recovery, and service-level objectives (SLOs).
You will also drive automation and operational efficiency by improving deployment processes, observability, monitoring, alerting, and infrastructure management. ๐ ๏ธ A key part of the role will be creating a culture of reliability, continuous improvement, knowledge sharing, and effective technical collaboration.
The ideal candidate should combine strong technical judgment with excellent leadership, communication, and problem-solving abilities. ๐ง You should be comfortable making data-driven decisions, managing operational priorities, and guiding teams through complex reliability challenges.
๐ฏ Qualifications
-
๐ Bachelorโs degree or equivalent qualification in Computer Science, Software Engineering, Information Technology, or a related field.
-
โ๏ธ Strong understanding of Site Reliability Engineering (SRE), DevOps, cloud infrastructure, and distributed systems.
-
โ๏ธ Familiarity with cloud platforms such as AWS, Azure, or Google Cloud.
-
๐ณ Knowledge of containers, Kubernetes, CI/CD pipelines, infrastructure as code, and automation tools.
-
๐ Strong understanding of monitoring, observability, logging, alerting, system performance, and service reliability metrics.
-
๐จ Knowledge of incident management, root cause analysis, disaster recovery, and capacity planning.
-
๐ Understanding of security, scalability, availability, and resilient system design principles.
-
๐ง Strong technical problem-solving and analytical skills with a focus on continuous improvement.
-
๐ค Excellent leadership, communication, mentoring, and cross-functional collaboration skills.
-
๐ Ability to define and monitor SLOs, SLIs, operational metrics, and reliability objectives.
-
๐ Strong documentation and decision-making skills with a structured approach to complex technical challenges.
-
๐ก Proactive mindset with a passion for automation, operational excellence, system reliability, and engineering innovation.