
Expert Site Reliability Engineer
TAWANTECH
Visa & sponsorship
- Employers in Saudi Arabia sponsor the residence visa by default, and nothing in the posting says otherwise.
Job description
Purpose:
To drive the reliability, availability, scalability, and operational resilience of critical technology services by applying advanced software engineering, automation, observability, and reliability engineering practices.
Main Duties and Responsibilities:
-
Define and implement advanced reliability engineering practices across critical technology services.
-
Establish and monitor Service Level Indicators (SLIs), Service Level Objectives (SLOs), and reliability targets.
-
Design automation to reduce manual operational activities and improve system resilience.
-
Develop and enhance monitoring, observability, alerting, and incident detection capabilities.
-
Lead technical analysis and resolution of complex production incidents.
-
Conduct root-cause analysis and drive permanent corrective and preventive actions.
-
Design solutions to improve system availability, scalability, capacity, and disaster resilience.
-
Identify reliability risks and recommend architectural and engineering improvements.
-
Drive performance engineering and capacity planning for critical services.
-
Provide advanced technical guidance and mentorship on SRE practices.
-
Promote automation and engineering approaches that reduce operational toil and improve service reliability.
Requirements
QUALIFICATIONS & REQUIREMENTS
-
Bachelor's degree in Computer Science, Software Engineering, IT, or a related field.
-
5+ years of experience in Site Reliability Engineering, DevOps, Platform Engineering, or related roles.
-
Strong experience in cloud platforms, Kubernetes, and production environments.
-
Strong knowledge of monitoring, observability, alerting, SLIs, SLOs, and reliability metrics.
-
Hands-on experience with automation, scripting, CI/CD, and Infrastructure as Code (Terraform).
-
Proven experience in complex incident management, troubleshooting, and Root Cause Analysis (RCA).
-
Strong understanding of high availability, scalability, performance engineering, capacity planning, and disaster recovery.
-
Experience driving reliability improvements and reducing operational toil through automation.
-
Strong analytical, problem-solving, and technical leadership skills.
-
Experience in Banking, FinTech, or Payment environments is preferred