
Site Reliability Engineer
Lucidya
Visa & sponsorship
- Employers in Saudi Arabia sponsor the residence visa by default, and nothing in the posting says otherwise.
Job description
About Lucidya
Lucidya is an AI-native platform for customer experience (CX) intelligence that manages entire customer lifecycles autonomously, from initial engagement through retention and growth.
Unlike platforms that only surface insights and leave the action to you, Lucidya closes the loop with proprietary NLU technology built in-house and trained on millions of multilingual conversations. This enables marketing, support, CX, and research teams to deliver personalized experiences that drive measurable improvements in customer satisfaction, retention, and lifetime value.
As we continue scaling globally, the reliability, performance, and resilience of our infrastructure become mission-critical to everything we do.
Why this role matters
At Lucidya, our platform processes massive volumes of real-time customer data. Any downtime, latency, or instability directly impacts our customersâ ability to make decisions and serve their own users.
This role exists to make sure that doesnât happen.
As a Site Reliability Engineer, youâll sit at the heart of our platformâs stability, owning the reliability of our cloud infrastructure and ensuring it scales seamlessly as we grow. You wonât just react to issues; youâll anticipate them, design systems that prevent them, and build automation that removes them entirely.
If you enjoy solving complex infrastructure challenges, eliminating inefficiencies, and building systems that âjust workâ - this is where youâll thrive.
What Youâll Do
Youâll be responsible for outcomes, not just tasks. Hereâs what success looks like in this role:
Youâll make reliability the default
-
Youâll design and maintain infrastructure that is highly available, fault-tolerant, and scalable
-
Youâll proactively identify and eliminate single points of failure before they become incidents
-
Youâll ensure our production systems remain stable, even under increasing scale and load
Youâll own and optimize our cloud environments
-
Youâll manage and continuously improve workloads across AWS, GCP, or Azure
-
Youâll use Infrastructure as Code (Terraform) to standardize and scale infrastructure
-
Youâll optimize resource usage to balance performance and cost
Youâll run and improve Kubernetes in production
-
Youâll operate and scale Kubernetes clusters (EKS, GKE, etc.) with confidence
-
Youâll troubleshoot issues quickly and ensure smooth deployments and upgrades
-
Youâll ensure our containerized workloads perform reliably at scale
Youâll build strong observability and respond to incidents
-
Youâll implement and refine monitoring systems using tools like Prometheus, Grafana, Datadog, or ELK
-
Youâll define alerting that is meaningful, not noisy
-
Youâll respond to incidents, lead root cause analysis, and ensure we learn from every failure
Youâll automate everything that shouldnât be manual
-
Youâll write scripts and build tooling to eliminate repetitive operational work
-
Youâll continuously improve infrastructure efficiency through automation
-
Youâll promote a culture where manual work is a temporary state, not the norm
Youâll collaborate to improve the entire system
-
Youâll work closely with DevOps and engineering teams to solve performance bottlenecks
-
Youâll contribute to CI/CD improvements and deployment reliability
-
Youâll help shape reliability best practices across the organization
What success looks like (First 90 Days)
First 30 days:
-
Youâve built a strong understanding of our infrastructure, systems, and workflows
-
Youâre contributing to day-to-day operations with support from the team
-
Youâve started identifying areas for improvement in automation and reliability
By 90 days:
-
Youâre independently managing infrastructure tasks and troubleshooting issues
-
Youâre actively contributing to reliability and scalability improvements
-
Youâve taken ownership of parts of our infrastructure and are improving them
Requirements
Who You Are
This is what will make you successful in this role:
-
Youâve spent ~3 years working in SRE, DevOps, or infrastructure engineering, and youâve seen what breaks at scale
-
Youâre comfortable working in cloud environments like AWS, GCP, or Azureâand you understand how distributed systems behave
-
Youâve worked hands-on with Kubernetes in production and know how to troubleshoot it when things go wrong
-
You donât just fix issues - you ask why they happened and make sure they donât happen again
Technically, you likely:
-
Use Terraform (or similar IaC tools) to manage infrastructure
-
Work confidently with Docker and Kubernetes
-
Write scripts in Python, Bash, or similar to automate workflows
-
Understand CI/CD pipelines (Jenkins, GitHub Actions, Bitbucket, etc.)
-
Have a solid grasp of networking, load balancing, and high-availability design
When it comes to monitoring:
-
Youâve implemented tools like Prometheus, Grafana, Datadog, or ELK
-
You know the difference between useful alerts and noise
-
You focus on signals that actually drive action
What sets you apart:
-
You take ownership - you donât wait to be told something is broken
-
Youâre calm under pressure and methodical during incidents
-
You simplify complexity instead of adding to it
-
You communicate clearly, even when explaining deeply technical issues
-
You care about building systems that make other engineers more effective
Nice to Have (but not required)
-
Experience with RabbitMQ or Redis in production
-
Familiarity with Ansible or AWX
-
Exposure to multi-cloud or hybrid environments
-
Cloud certifications (AWS, GCP) or Linux certifications
-
Background from ITI (Information Technology Institute)
What the hiring process will look like
-
Screening Interview â Talent Acquisition
-
Technical Interview â SRE Lead
-
Technical Task
-
Final Interview â SRE Lead & Cloud DevOps Director