
Senior Site Reliability Engineer
Lloyds
Job description
-
As a Senior Site Reliability Engineer, you’ll be an experienced engineering practitioner responsible for improving the reliability, scalability and operational excellence of cloud-hosted services!
-
Lead the operation, reliability, scalability and performance improvement of cloud-hosted applications and services
-
Drive automation initiatives that reduce manual effort, improve consistency and strengthen service resilience
-
Partner with engineering, platform and product teams to deliver and continuously improve cloud solutions and security data platforms
-
Lead incident response, problem investigations and root cause analysis activities, ensuring learning is converted into measurable improvements
-
Develop and enhance observability, monitoring and alerting capabilities that provide actionable insight into service health and customer impact
-
Mentor engineers and contribute to engineering standards, patterns and communities of practice across the organisation
Benefits
-
A generous holiday allowance: You’ll be eligible for a minimum of 22 days holiday (excluding bank holidays), rising to 30 days based on length of service and grade.
-
A flexible way of working: Whether you want flexibility over your location or when you log on, together we can create an approach that works for you and for the business.
-
Family leave: Up to 63 weeks of maternity or adoption leave. Statutory maternity or adoption pay is available for 39 weeks, and 20 weeks will be enhanced to the equivalent of full pay. Partners can have six weeks of fully paid paternity leave.
-
Flex cash: This is 4% of your basic salary and can be used to spend on the benefits of your choice, or you can choose to take it as a cash top up in your monthly salary.
-
Health insurance: Our company funded Private Medical Benefit provides all colleagues with access to good quality medical care, including accommodation, nursing care and specialist advice.
-
Colleague Offers: Get discounts on everything from electrical items to cinema tickets and weekly food shopping. You can share this benefit with up to ten family members or friends.
-
Financial products: Take advantage of our great financial products, some at a discounted rate, including current accounts, home and car insurance and loans.
-
Share plans: Participate in Sharematch and receive matching shares of up to £45 a month from the company, and you can choose to participate in Sharesave, our combined savings and share option plan.
-
Pension: We offer a generous pension plan, with all joiners being automatically enrolled in our ‘Your Tomorrow’ scheme. You can decide how much you save and get a say in where your contributions are invested.- Strong scripting or programming capability using languages such as Python, PowerShell or Bash to automate operational activities and improve reliability outcomes
-
We’re looking for experienced engineers who are passionate about reliability engineering, cloud technologies and operational excellence
-
Problem Solving and Critical Thinking skills, with experience diagnosing complex production incidents and driving effective remediation. These are recognised enterprise-wide skills in the SBO Skills Library
-
DevOps expertise, including automation, continuous integration, continuous delivery and collaborative software delivery practices. DevOps is identified as a core engineering skill within the enterprise skills taxonomy
-
Collaboration and Impactful Communication skills, with the ability to build strong working relationships and explain technical concepts to both technical and non-technical audiences. These are recognised enterprise-wide skills in the SBO Skills Library
-
Strong experience working with public cloud platforms, particularly Google Cloud Platform (GCP), within operational or engineering environments
-
Experience implementing and maintaining Infrastructure as Code solutions and engineering platforms using tools such as Terraform
-
SRE & Service Engineering skills, with significant experience improving the reliability, availability and supportability of production services. The SBO Skills Library identifies SRE & Service Engineering as a core engineering capability
-
Experience with Google SecOps
-
Experience working in Site Reliability Engineering, Platform Engineering, Software Engineering or DevOps roles supporting complex production environments
-
Experience with monitoring and observability platforms such as Dynatrace
-
Experience working with Jira and Confluence within Agile delivery teams
-
Experience using Terraform, GitHub, Kubernetes or similar engineering tooling
-
Professional certifications in GCP or comparable cloud technologies
-
Experience coaching engineers, developing reusable engineering patterns and improving operational standards
-
Interest / passion in AI and the application of agentic solutions would be advantageous. This lab is at the forefront of agentic use cases and AI security, and whilst professional experience may be limited, personal interest & development would be great- End Date: Wednesday 09 September 2026