Soar logo

Site Reliability Engineer Manager (SREM)

Soar

On-site๐Ÿ‡ธ๐Ÿ‡ฆRiyadh, 01, Saudi ArabiaseniorPosted 1d ago

Visa & sponsorship

  • Employers in Saudi Arabia sponsor the residence visa by default, and nothing in the posting says otherwise.

Job description

About The Role

Role Summary:

We are looking for a Senior Site Reliability Engineer (SRE) to champion reliability, performance, and efficiency across our engineering organization. In this role, you will bridge the gap between development and operations by applying software engineering principles to system administration. You will be instrumental in building fault-tolerant infrastructure, establishing SLOs, and ensuring our platforms remain highly available and secure, adhering to strict financial and data compliance standards in KSA.

Key Responsibilities:

  • Drive Reliability: Define, measure, and manage Service Level Indicators (SLIs), Service Level Objectives (SLOs), and Error Budgets to balance feature velocity with system stability.

  • Architect Infrastructure: Design, build, and maintain our cloud-native infrastructure on [AWS / GCP / Azure] using Infrastructure as Code (IaC) tools like [Terraform / Pulumi].

  • Automate Everything: Ruthlessly eliminate manual toil by building automation tools for provisioning, scaling, failover, and deployment.

  • Manage Incidents: Lead the incident management process, participate in on-call rotations, and conduct blameless post-mortems to ensure continuous improvement and prevent recurring issues.

  • Enhance Observability: Implement comprehensive monitoring, logging, and alerting systems using tools like [Prometheus, Grafana, Datadog, or ELK stack] to proactively identify system bottlenecks.

  • Optimize CI/CD: Partner with development teams to optimize deployment pipelines for speed, safety, and reliability.

  • Ensure Compliance & Security: Collaborate with the security team to implement robust security practices and ensure infrastructure complies with KSA regulations (e.g., SAMA, NCA, and PDPL data localization requirements).

  • Mentorship: Mentor junior engineers and cultivate a strong culture of reliability, operational excellence, and shared ownership across the engineering team.

What We Are Looking For:

  • Experience: 5+ years in Software Engineering, DevOps, or SRE roles, with at least 2 years in a senior capacity.

  • Cloud Expertise: Deep, hands-on experience with major cloud providers, preferably [AWS / GCP].

  • Containerization: Strong proficiency in container orchestration, specifically Kubernetes and Docker.

  • Coding Skills: Strong programming skills in at least one language such as Go, Python, or Java (not just bash scripting).

  • Infrastructure as Code: Proven experience with Terraform, Ansible, or similar IaC tools.

  • Systems Knowledge: Deep understanding of Linux operating systems, networking (TCP/IP, DNS, routing), and distributed systems architecture.

  • Communication: Excellent verbal and written communication skills in English (Arabic is a strong plus).

Nice to Have:

  • Prior experience working in a Fintech, Proptech, or highly regulated environment.

  • Familiarity with Saudi Arabian financial and data privacy regulations (SAMA, NCA frameworks).

  • Experience with database reliability (PostgreSQL, Redis, Kafka).