Jobgether logo

Site Reliability Engineering Team Lead (Principal SRE)

Jobgether

RemoteleadPosted 1h ago

Job description

This position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for a Site Reliability Engineering Team Lead (Principal SRE) based in the United States.

This Principal-level role owns the reliability, availability, and operational health of a global cloud-native AI platform.

You will combine deep hands-on SRE expertise with technical leadership, shaping reliability strategy across critical production services.

The role encompasses SLI/SLO/SLA governance, observability, automation, incident response, production readiness, and high-risk change management.

You will work closely with engineering, DevOps, platform, architecture, and operations teams to embed reliability throughout the software development lifecycle.

As a technical leader without direct reports, you will influence through expertise, mentorship, standards, and informed decision-making.

The environment is distributed and highly technical, with complex systems requiring strong availability, scalability, and operational resilience.

This opportunity is ideal for an experienced SRE professional who enjoys solving challenging infrastructure problems while building sustainable reliability practices.

Accountabilities

Provide technical leadership across the Site Reliability Engineering function, helping select, mentor, and develop engineers across multiple locations.

Establish technical direction, priorities, engineering standards, and reliability practices while contributing performance and growth feedback to team managers.

Own and execute a reliability roadmap covering a 2–3 quarter planning horizon.

Define and govern SLI, SLO, and SLA frameworks supporting contracted availability targets of up to 99.95%.

Design and maintain a sustainable on-call model while monitoring operational workload, page volume, and team health.

Serve as a Tier 2 technical escalation point for major production incidents and collaborate with incident management and operations teams.

Promote a blameless postmortem culture and ensure incident reviews result in actionable systemic improvements.

Lead Production Readiness and non-functional requirements reviews with development teams.

Contribute to root cause analysis and drive reliability improvements resulting from production incidents.

Act as an approval authority for high-risk and out-of-window production changes.

Define strategic direction for metrics, dashboards, alerting, SLI/SLO monitoring, escalation, and automation.

Drive CI/CD automation for service deployments, rollbacks, and operational processes.

Partner with DevOps and platform teams to evolve shared infrastructure and reliability capabilities.

Work with engineering managers and architects to incorporate reliability principles into the SDLC by default.

Participate in architecture reviews and reliability consulting while clearly communicating technical risks and reliability posture to technical and non-technical stakeholders.

Requirements

8+ years of hands-on experience in Site Reliability Engineering, DevOps, cloud platforms, or closely related roles, including experience leading a team or owning a technical function.

Demonstrated ability to establish technical direction, maintain engineering standards, and influence teams through technical authority, with or without formal management responsibility.

Hands-on experience with container orchestration and service technologies such as Kubernetes, Docker, and Istio.

Strong experience with public cloud platforms, particularly Azure, with exposure to AWS and Google Cloud.

Experience with observability technologies covering metrics, dashboards, and alerting, such as Zabbix, Prometheus, and Grafana.

Experience designing and operating CI/CD pipelines and infrastructure-as-code solutions, including technologies such as Terraform and Flux.

Proficiency in at least one scripting or programming language, such as Python, Go, or Shell.

Strong UNIX/Linux expertise, including system configuration, performance troubleshooting, and networking fundamentals such as Layer 4/5, DNS, HTTP/S, and TLS.

Strong understanding of high-availability architecture, including redundancy, failover strategies, and blast-radius management.

Excellent written and verbal communication skills in English, with the ability to explain complex technical concepts clearly.

Previous SRE leadership experience and experience managing or influencing distributed technical teams are preferred.

Experience with log aggregation and analytics platforms such as Loki or Thanos is preferred.

Familiarity with ITSM and project management tools such as Jira and Confluence is beneficial.

Experience in automotive, embedded systems, or other latency-sensitive production environments is advantageous.

Strong collaborative mindset, sound judgment under pressure, and the ability to operate effectively in ambiguous and technically complex environments.

Benefits

Competitive compensation and benefits package.

Annual bonus opportunity.

Medical, dental, and vision insurance coverage.

Life and disability insurance.

Paid time off and paid holidays.

Company contribution to an RRSP retirement savings plan.

Equity awards for eligible positions and levels.

Remote and/or hybrid work options depending on the position and location.

Opportunity to work on large-scale cloud-native AI and connected technology platforms.

Exposure to distributed engineering teams and complex global production environments.

Opportunities to influence technical strategy, reliability standards, and engineering practices at a Principal level.

A collaborative environment focused on innovation, technical growth, and continuous improvement.

How Jobgether works:

We use an AI-powered matching process to ensure your application is reviewed quickly, objectively, and fairly against the role's core requirements. Our system identifies the top-fitting candidates, and this shortlist is then shared directly with the hiring company. The final decision and next steps (interviews, assessments) are managed by their internal team.

We appreciate your interest and wish you the best!

 Why Apply Through Jobgether?

Data Privacy Notice: By submitting your application, you acknowledge that Jobgether will process your personal data to evaluate your candidacy and share relevant information with the hiring employer. This processing is based on legitimate interest and pre-contractual measures under applicable data protection laws (including GDPR). You may exercise your rights (access, rectification, erasure, objection) at any time.

#LI-CL1