SmartChoice International GCC logo

Senior DevOps Engineer

SmartChoice International GCC

On-site๐Ÿ‡ธ๐Ÿ‡ฆRiyadh, 01, Saudi ArabiaseniorPosted 1d ago

Visa & sponsorship

  • Employers in Saudi Arabia sponsor the residence visa by default, and nothing in the posting says otherwise.

Job description

Senior DevOps Engineer

Full-time | Riyadh

We are partnering with a growing technology organisation to appoint a Senior DevOps Engineer to build and operate the infrastructure behind a rapidly scaling technology platform.

This is a hands-on role spanning cloud infrastructure, automation, CI/CD, container platforms, security, observability and production reliability. You will work closely with engineering teams to create scalable, secure and resilient environments, while establishing the operational standards the wider technology function can build on.

Key Responsibilities

  • Design and maintain scalable, reproducible infrastructure using infrastructure as code.

  • Own CI/CD pipelines covering build, testing, security and production deployment.

  • Manage containerised workloads across Kubernetes and cloud-based compute environments.

  • Build safe deployment strategies with appropriate rollout, rollback and recovery mechanisms.

  • Manage cloud networking, API gateways, load balancing, DNS, certificates and service connectivity.

  • Implement identity, secrets management and least-privilege access across environments.

  • Operate production databases, including high availability, replication, backup and recovery.

  • Establish monitoring, alerting and operational metrics across infrastructure and applications.

  • Lead incident response and continuously improve reliability, security and operational resilience.

  • Optimise infrastructure utilisation and costs, including high-compute workloads.

  • Help create reusable platform capabilities that enable engineering teams to deploy and operate services independently.

What We're Looking For

  • 8+ years of experience in DevOps, SRE, infrastructure or platform engineering.

  • Strong hands-on experience with

    AWS

    , including compute, networking, IAM and storage.

  • Strong infrastructure-as-code experience with

    Terraform or OpenTofu

    , including reusable modules and state management.

  • Experience with

    Ansible

    or comparable configuration management.

  • Strong CI/CD experience, particularly

    GitHub Actions

    or equivalent.

  • Strong experience with

    Docker, Kubernetes and Helm

    , with EKS highly desirable.

  • Experience with

    GitOps

    , including ArgoCD or Flux.

  • Experience with API gateways, authentication, routing and rate limiting.

  • Production experience with

    PostgreSQL, RDS/Aurora and Redis

    or equivalent technologies.

  • Experience with

    Vault, AWS Secrets Manager or Parameter Store

    .

  • Strong observability experience with

    Prometheus, Grafana, CloudWatch and OpenTelemetry

    or equivalents.

  • Comfortable working in Linux environments with

    Bash and Python

    .

  • Strong understanding of security, reliability, incident management and production operations.

Advanced Infrastructure

Experience operating AI, machine learning or other high-compute workloads in production is highly desirable.

Relevant experience may include:

  • GPU infrastructure and workload scheduling.

  • Model-serving or inference platforms.

  • Autoscaling and capacity management for compute-intensive workloads.

  • Managing performance, utilisation and cost of high-compute environments.

  • Technologies such as

    vLLM, Triton, TGI, KServe, Ray Serve, Amazon Bedrock or SageMaker

    .

Technical Environment

AWS | Terraform/OpenTofu | Ansible | GitHub Actions | Docker | Kubernetes/EKS | Helm | ArgoCD/Flux | API Gateways | PostgreSQL | RDS/Aurora | Redis | Vault/Secrets Manager | Prometheus | Grafana | CloudWatch | OpenTelemetry | Linux | Bash | Python

We're looking for an engineer with strong technical depth and operational judgement who can work calmly through complex production problems, understand the impact of infrastructure changes, and build systems with reliability, security and cost in mind.