R

Site Reliability Engineer (Production Reliability, Azure Operations & Databricks)

Realign

Hybrid๐Ÿ‡จ๐Ÿ‡ฆToronto, ON, CanadamidPosted 1d ago

Job description

Toronto, Ontario M5V 3L9 Posted September 3rd, 2026

Job Type: Full Time
Job Category: IT

Job Description

Job Title: Site Reliability Engineer (Production Reliability, Azure Operations & Databricks)

Job Type: Full-Time

Location: Toronto, ON (Hybrid)

Job Overview

As an Intermediate Site Reliability Engineer, you will maintain, optimize, and ensure the production reliability of enterprise-level Azure and Databricks platforms. You will focus on platform availability, continuous monitoring, incident response, operational readiness, and platform support while collaborating closely with engineering, security, network, and data teams.

Key Responsibilities

  • Platform Reliability & Support: Monitor and support production Azure and Databricks environments to ensure maximum availability, performance, and operational readiness.

  • Incident Response & On-Call: Respond to production incidents, participate in on-call support rotations, join incident bridges, and execute emergency changes when required.

  • Databricks & Data Services: Support Databricks workspaces, clusters, workflows, user access, and Unity Catalog governance (catalogs, schemas, storage credentials). Maintain integrations with ADLS Gen2, Azure Data Factory, Azure SQL, and Key Vault.

  • Infrastructure & Networking: Troubleshoot Azure infrastructure, including storage accounts, Blob Storage, VNets, NSGs, private endpoints, DNS, and hub-and-spoke connectivity.

  • Observability & Monitoring: Manage alerts, dashboards, and platform health using Azure Monitor, Log Analytics, Grafana, Prometheus, Dynatrace, Datadog, or New Relic.

  • Root Cause & Maintenance: Perform root cause analysis (RCA), problem management, system patching, upgrades, maintenance, and disaster recovery exercises.

  • Operations & Documentation: Maintain operational runbooks and knowledge articles while tracking tickets and tasks in JIRA and ServiceNow.

Required Technical Qualifications (Must-Haves)

  • Experience: 3+ years supporting Azure production cloud infrastructure and 1+ years supporting Databricks environments.

  • OS Administration: 1+ years of Windows Server administration and 1+ years of Linux administration.

  • Data & Storage: Proven experience supporting Azure Storage services, including ADLS Gen2 and Blob Storage.

  • Networking & Security: Understanding of VNets, NSGs, private endpoints, DNS, routing, Entra ID (Azure AD), RBAC, managed identities, and Azure Key Vault.

  • Monitoring Tools: Hands-on experience with Azure Monitor, Log Analytics, Grafana, Prometheus, Dynatrace, Datadog, or New Relic.

  • ITSM & Operations: Hands-on experience with incident escalation, change management, RCA, JIRA, ServiceNow, and operational runbooks.

Preferred Qualifications (Nice-to-Haves)

  • Operational support experience with Azure SQL and Azure Data Factory (integration runtimes, linked services, orchestration).

  • Understanding of Unity Catalog governance and Disaster Recovery/Business Continuity (RTO/RPO).

  • Exposure to AI/GenAI platforms, Azure OpenAI, MLOps, model endpoints, or RAG services.

  • Experience in cost monitoring, capacity planning, and platform health reporting.

Key Competencies & Soft Skills

  • Strong analytical and calm problem-solving mindset during critical production incidents.

  • Excellent cross-team collaboration skills (working with network, security, and platform teams).

  • Strong documentation skills and customer-focused approach to platform reliability.

Required Skills

Cloud Developer Data Analyst DevOps Engineer Senior Email Security Engineer