
Lead Engineer โ Infrastructure & Cloud Engineering
SmartChoice International GCC
Visa & sponsorship
- Employers in UAE sponsor the residence visa by default, and nothing in the posting says otherwise.
Job description
Lead Engineer โ Infrastructure & Cloud Engineering
Abu Dhabi, UAE | Permanent
About the Client
Our client is a high-growth technology powerhouse operating at the forefront of AI, cloud and digital infrastructure. With an ambitious vision for the future of technology, the organization is bringing together world-class engineering talent, advanced infrastructure and next-generation platforms to solve some of the most complex technology challenges at scale.
This is an opportunity to join an environment where deep engineering, ambitious technology and genuine innovationcome together, with the chance to influence the architecture of critical platforms and help shape what comes next.
The Opportunity
We are looking for an experienced
Lead Engineer โ Infrastructure & Cloud Engineering
to provide technical leadership across private cloud, virtualization, container and observability platforms.
The role combines deep hands-on engineering expertise with technical leadership, helping shape platform strategy, engineering standards, operational excellence and the adoption of emerging technologies.
You will work closely with architecture, security, platform engineering, SRE and operations teams to build, operate and continuously improve highly available infrastructure platforms supporting demanding technology environments.
This is a hands-on technical leadership role for someone who enjoys solving complex infrastructure challenges, leading engineering initiatives and helping teams adopt modern approaches to cloud, automation and platform engineering.
Key Responsibilities
-
Lead the technical design, implementation and lifecycle management of large-scale
private cloud, virtualization and container platforms
, including OpenStack and OpenShift or comparable technologies.
-
Define and drive engineering standards, design principles, operational best practices and technical governance across infrastructure and cloud platforms.
-
Provide technical leadership and mentorship to engineering teams, supporting technical decision-making, complex problem-solving, knowledge sharing and continuous development.
-
Lead the strategy, design and adoption of
observability platforms
, covering metrics, logs, traces, dashboards, alerting, service health and SLO/SLI capabilities.
-
Drive the adoption of
AI-assisted operations and intelligent automation
to improve operational efficiency, incident management, root cause analysis, platform reliability and service automation.
-
Lead the evaluation, integration and adoption of new technologies, products and architectural approaches in collaboration with Architecture, Product Engineering, SRE, Security and Operations teams.
-
Provide technical oversight for complex platform upgrades, migrations, production changes and infrastructure transformation initiatives.
-
Act as a senior technical escalation point for critical production incidents, leading complex troubleshooting, incident response, root cause analysis and service recovery.
-
Drive capacity planning, scalability, performance optimisation, resilience engineering and operational readiness across infrastructure platforms.
-
Ensure observability, automation and operational tooling are effectively integrated with service management and incident management processes.
-
Work closely with security, compliance and risk teams to ensure infrastructure platforms meet appropriate cybersecurity, governance and regulatory requirements.
-
Define and champion automation strategies using
Infrastructure as Code, GitOps, CI/CD and platform engineering
practices.
-
Lead the development and maintenance of technical standards, architecture documentation, operational procedures, design guides and knowledge resources.
-
Collaborate with leadership teams on infrastructure strategy, roadmap planning, technology evaluation and long-term platform evolution.
What We're Looking For
-
Bachelor's or Master's degree in Computer Science, Engineering, Software Engineering or a related technology discipline, or equivalent practical experience.
-
8+ years of experience
designing, implementing, operating, troubleshooting and leading large-scale cloud, infrastructure or platform engineering environments.
-
Strong hands-on experience with
private cloud and virtualization platforms
, with proven experience in OpenStack, OpenShift or comparable technologies.
-
Experience working with
container platforms and orchestration technologies
, with a good understanding of containerised infrastructure.
-
Proven experience providing technical leadership across complex infrastructure engineering environments.
-
Strong understanding of compute infrastructure, including
x86/ARM architecture, Linux, KVM, virtualization, server hardware, firmware and infrastructure lifecycle management
.
-
Expert-level Linux administration and troubleshooting skills, including performance analysis, system tuning, patching and operational support.
-
Strong understanding of data centre networking, including
TCP/IP, routing, firewalls, load balancing, VLAN/VXLAN, DNS, DHCP and related technologies
.
-
Deep understanding of observability principles, including
metrics, logs, traces, alerting, dashboards and service-level monitoring
.
-
Hands-on experience with modern observability technologies such as
Prometheus, Grafana, OpenTelemetry, ELK/OpenSearch or comparable platforms
.
-
Proven experience operating large-scale private or public cloud environments, highly available infrastructure or mission-critical technology platforms.
-
Strong experience with
Infrastructure as Code, automation, CI/CD and GitOps
, using technologies such as Terraform, Ansible, Helm or comparable tools.
-
Strong scripting or programming capabilities using
Python, Go, Bash or similar languages
.
-
Experience with
AI-assisted operations, workflow automation or intelligent infrastructure management
is highly desirable.
-
Strong understanding of software-defined infrastructure, platform engineering, reliability engineering, incident management and infrastructure lifecycle automation.
-
Experience integrating observability and operational tooling with service management and incident management processes.
-
Understanding of infrastructure security, security monitoring, compliance, data governance and risk management.
-
Relevant certifications across cloud, Linux, virtualization, Kubernetes, OpenStack, OpenShift or IT service management are advantageous.
-
Excellent analytical, troubleshooting, communication, documentation, stakeholder management, mentoring and technical leadership skills.
You'll be a technically strong infrastructure engineer who is comfortable operating at both hands-on engineering and technical leadership level.
You'll be able to get into the detail of complex infrastructure problems while also stepping back to define standards, influence technical decisions and guide engineering teams.
You should have a strong interest in cloud infrastructure, automation, observability and emerging technology, with the ability to identify where new approaches โ including AI-assisted operations โ can genuinely improve reliability and operational efficiency.