TAWANTECH logo

Expert Site Reliability Engineer

TAWANTECH

On-site๐Ÿ‡ธ๐Ÿ‡ฆRiyadh, 01, Saudi ArabiaseniorPosted 4d ago

Visa & sponsorship

  • Employers in Saudi Arabia sponsor the residence visa by default, and nothing in the posting says otherwise.

Job description

Purpose:

To drive the reliability, availability, scalability, and operational resilience of critical technology services by applying advanced software engineering, automation, observability, and reliability engineering practices.

Main Duties and Responsibilities:

  • Define and implement advanced reliability engineering practices across critical technology services.

  • Establish and monitor Service Level Indicators (SLIs), Service Level Objectives (SLOs), and reliability targets.

  • Design automation to reduce manual operational activities and improve system resilience.

  • Develop and enhance monitoring, observability, alerting, and incident detection capabilities.

  • Lead technical analysis and resolution of complex production incidents.

  • Conduct root-cause analysis and drive permanent corrective and preventive actions.

  • Design solutions to improve system availability, scalability, capacity, and disaster resilience.

  • Identify reliability risks and recommend architectural and engineering improvements.

  • Drive performance engineering and capacity planning for critical services.

  • Provide advanced technical guidance and mentorship on SRE practices.

  • Promote automation and engineering approaches that reduce operational toil and improve service reliability.

Requirements

QUALIFICATIONS & REQUIREMENTS

  • Bachelor's degree in Computer Science, Software Engineering, IT, or a related field.

  • 5+ years of experience in Site Reliability Engineering, DevOps, Platform Engineering, or related roles.

  • Strong experience in cloud platforms, Kubernetes, and production environments.

  • Strong knowledge of monitoring, observability, alerting, SLIs, SLOs, and reliability metrics.

  • Hands-on experience with automation, scripting, CI/CD, and Infrastructure as Code (Terraform).

  • Proven experience in complex incident management, troubleshooting, and Root Cause Analysis (RCA).

  • Strong understanding of high availability, scalability, performance engineering, capacity planning, and disaster recovery.

  • Experience driving reliability improvements and reducing operational toil through automation.

  • Strong analytical, problem-solving, and technical leadership skills.

  • Experience in Banking, FinTech, or Payment environments is preferred