Throne Solutions logo

Site Reliability Engineer

Throne Solutions

On-site๐Ÿ‡ธ๐Ÿ‡ฆRiyadh, 01, Saudi ArabiaseniorSAR 15,000โ€“17,000/moPosted 15d ago

Visa & sponsorship

  • Employers in Saudi Arabia sponsor the residence visa by default, and nothing in the posting says otherwise.

Job description

Site Reliability Engineer (SRE)

Company:

Throne Solutions

Location:

Riyadh, Kingdom of Saudi Arabia (KSA)

Employment Type:

Full-Time Freelance Contract

Work Arrangement:

Onsite

Experience:

5โ€“8 Years

Number of Positions:

2

Salary / Rate:

SAR 15,000 โ€“ 17,000 per month

Job Summary

Throne Solutions is seeking experienced

Site Reliability Engineers (SRE)

to join our technical team in

Riyadh, KSA

on a full-time freelance contract.

The SRE will be responsible for ensuring the

availability, reliability, scalability, performance, and security

of enterprise applications and infrastructure. The role combines systems engineering, cloud operations, automation, monitoring, incident management, and DevOps practices.

The ideal candidate should have strong hands-on experience with

Linux/Windows systems, cloud platforms, Kubernetes, containers, CI/CD, monitoring, automation, and production operations

.

Key ResponsibilitiesSite Reliability & Production Operations

  • Ensure high availability, reliability, and performance of production systems.

  • Monitor infrastructure and application health across production environments.

  • Identify and resolve performance, availability, and capacity issues.

  • Participate in incident, problem, and change management processes.

  • Provide L2/L3 support for critical production incidents.

  • Perform root-cause analysis (RCA) and implement permanent corrective actions.

  • Participate in planned maintenance, upgrades, and production deployments.

  • Support 24x7 operations and on-call activities when required.

Monitoring & Observability

  • Implement and maintain comprehensive monitoring and alerting solutions.

  • Monitor system availability, CPU, memory, storage, network, application performance, and service health.

  • Develop dashboards and meaningful alerts for proactive issue detection.

  • Work with tools such as Prometheus, Grafana, ELK/Elastic Stack, Splunk, Datadog, Nagios, or Zabbix.

  • Define and monitor

    SLIs, SLOs, and SLAs

    .

  • Reduce alert noise and improve incident detection and response.

Cloud Infrastructure

  • Administer and support cloud infrastructure across AWS, Microsoft Azure, or Google Cloud.

  • Manage compute, storage, networking, IAM, and cloud monitoring services.

  • Troubleshoot cloud infrastructure and connectivity issues.

  • Support cloud migration, optimization, and modernization initiatives.

  • Implement best practices for cloud availability, scalability, security, and cost optimization.

Kubernetes & Containers

  • Deploy, manage, and troubleshoot containerized applications.

  • Support Kubernetes clusters in production environments.

  • Troubleshoot pods, deployments, services, ingress, networking, and persistent volumes.

  • Work with Docker and Kubernetes technologies.

  • Monitor cluster health and resource utilization.

  • Support Kubernetes upgrades, scaling, and availability improvements.

Automation & Infrastructure as Code

  • Automate repetitive operational and infrastructure tasks.

  • Develop scripts and automation using Python, Bash, or PowerShell.

  • Use Infrastructure as Code tools such as Terraform and Ansible.

  • Automate infrastructure provisioning, configuration, deployment, and operational processes.

  • Improve operational efficiency through automation and self-healing mechanisms.

CI/CD & DevOps

  • Support and maintain CI/CD pipelines.

  • Work with Jenkins, GitLab CI/CD, GitHub Actions, Azure DevOps, or similar platforms.

  • Automate application and infrastructure deployment processes.

  • Collaborate with development teams to improve deployment reliability.

  • Implement rolling, blue-green, or canary deployment strategies where applicable.

  • Support version control and release-management processes.

Incident & Problem Management

  • Respond to production incidents within defined SLAs.

  • Participate in troubleshooting high-priority and critical incidents.

  • Perform detailed root-cause analysis.

  • Prepare incident reports and post-incident reviews.

  • Identify recurring issues and implement preventive measures.

  • Maintain operational runbooks and knowledge-base documentation.

  • Coordinate with application, network, security, database, and infrastructure teams during major incidents.

Required Technical Skills

  • Linux / Windows administration

  • AWS / Azure / GCP

  • Kubernetes

  • Docker / Containers

  • CI/CD

  • Git

  • Terraform / Ansible

  • Python / Bash / PowerShell

  • Prometheus / Grafana

  • ELK / Splunk or similar monitoring platforms

  • TCP/IP, DNS, HTTP/HTTPS, load balancing, and network troubleshooting

  • Incident and problem management

  • Infrastructure automation

  • Production support

Required Experience

  • 5โ€“8 years of experience

    in SRE, DevOps, Cloud Operations, Systems Engineering, or Production Engineering.

  • Strong experience supporting business-critical production environments.

  • Hands-on experience with cloud infrastructure and automation.

  • Experience troubleshooting complex production incidents.

  • Experience with monitoring, alerting, logging, and observability.

  • Practical experience with Kubernetes and containerized environments.

  • Experience with CI/CD pipelines and Infrastructure as Code.

  • Strong understanding of reliability, scalability, and high-availability concepts.

Preferred Certifications

  • AWS Certified Solutions Architect / DevOps Engineer

  • Microsoft Azure Administrator / Azure DevOps Engineer

  • Certified Kubernetes Administrator (CKA)

  • Certified Kubernetes Application Developer (CKAD)

  • HashiCorp Terraform Associate

  • Red Hat certifications

  • ITIL certification

Key Performance Indicators

  • Production availability and uptime.

  • SLA/SLO adherence.

  • Mean Time to Detect (MTTD).

  • Mean Time to Recovery/Resolve (MTTR).

  • Reduction in recurring incidents.

  • Successful automation and operational efficiency improvements.

  • Deployment success rate.

  • Monitoring and alerting effectiveness.

  • Quality and timeliness of RCA and post-incident documentation.

Candidate Profile

The ideal candidate should have:

  • Strong troubleshooting and analytical skills.

  • A proactive approach to reliability and automation.

  • Ability to work effectively during high-severity production incidents.

  • Strong communication and documentation skills.

  • Ability to collaborate with development, network, security, database, and infrastructure teams.

  • Ability to work independently in an onsite customer environment.

  • Willingness to participate in scheduled maintenance and on-call support when required.

Contract & Compensation Details

RequirementDetails

Job Title

Site Reliability Engineer (SRE)

Company

Throne Solutions

Location

Riyadh, KSA

Positions

2

Experience

5โ€“8 Years

Employment Type

Full-Time Freelance Contract

Work Arrangement

Onsite

Salary / RateSAR 15,000 โ€“ 17,000 per monthPrimary Skills

SRE, Cloud, Kubernetes, DevOps, Automation, Monitoring

Preferred Certifications

AWS / Azure / CKA / Terraform

Candidates whose salary expectations fall within SAR 15,000โ€“17,000 per month and who have strong production support, cloud, Kubernetes, automation, monitoring, and incident-management experience are encouraged to apply.