D

Senior Site Reliability Engineer (SRE)

Dental Intelligence

RemoteseniorPosted 3h ago

Job description

** About The Role: **

**

We're looking for a Senior Site Reliability Engineer to help us mature and scale the infrastructure behind our multi-cloud SaaS platform. Most of our footprint runs on Microsoft Azure, built from the ground up around cloud architecture principles: autoscaling App Service and Container Apps workloads, VM Scale Sets, and Kafka-based event streaming form the backbone. We also have a smaller, well-run presence on AWS, and a Windows-based on-premises component that sits at the core of our product and integrates with our cloud environment. This role is primarily focused on Azure, where the biggest opportunity for impact lives, and you'll work across the full multi-cloud picture. We'd like to see the on-prem and cloud sides operate as a more unified, well-integrated system than they do today, and that integration work is part of what makes this role interesting.

This is a high-impact role for someone who thinks in terms of systems rather than tickets, and who treats infrastructure like software. You'll have significant ownership over how our cloud infrastructure is architected, provisioned, secured, and operated going forward. If you get energized by taking a fast-growing environment and giving it real architectural rigor (consistent patterns, full infrastructure-as-code coverage, sane permission models, and cost discipline), this role was built for you.

We're looking for someone who wants to build the operating model, establish the standards, and lead the transformation. You should bring genuine software engineering discipline to how that infrastructure work gets done: everything in git, everything reviewed, everything automated. No snowflakes, no manual changes made "just this once."

Location: ** This is a fully remote role available to candidates located in U.S. states where Dental Intelligence currently has employees.

**

What You'll do: **

  • Own the reliability, scalability, and security posture of our Azure environment end-to-end

  • Lead the effort to bring our infrastructure fully under Terraform-managed IaC, replacing manual and ad-hoc provisioning with repeatable, version-controlled deployments

  • Treat infrastructure code like production software: everything lives in git, changes go through pull requests and peer review, and modules are tested before they ship

  • Define and implement a coherent Azure architecture strategy, including resource organization, naming and tagging standards, subscription and management group hierarchy, and network topology

  • Redesign and enforce least-privilege access and permission boundaries across Azure RBAC, Entra ID, and service principals

  • Identify and eliminate wasteful or redundant resource provisioning, and build cost visibility and accountability into how infrastructure is deployed

  • Build CI/CD pipelines for infrastructure changes so that plan/apply, validation, and policy checks are automated rather than run by hand

  • Build monitoring, alerting, and observability practices (Azure Monitor, Log Analytics, App Insights, or equivalent) that give the team real signal

  • Manage and modernize the Windows-based on-prem component that sits at the core of our product, and work to integrate it more tightly with our Azure environment

  • Bring the same IaC and automation discipline to bear on our AWS footprint as needed, keeping it as clean and well-run as it is today

  • Drive incident response, postmortems, and reliability engineering practices such as SLOs, error budgets, and capacity planning

  • Partner closely with engineering teams to bake reliability, security, and IaC discipline into the software delivery lifecycle

  • Mentor other engineers on cloud, IaC, and software engineering best practices, raising the bar across the team

**

What You Bring: **

  • โ€‹โ€‹6+ years in SRE, DevOps, or infrastructure engineering roles, with deep, hands-on Azure experience

  • Strong, demonstrable Terraform expertise is required. You should be comfortable designing module structures, managing state, and using Terraform to manage complex, multi-resource environments from scratch

  • A genuine cloud mentality: you default to automation, reproducibility, and IaC over manual changes, and you get uncomfortable when infrastructure can't be traced back to code

  • A real software engineering mindset applied to infrastructure: fluency with git workflows (branching, PRs, code review), a bias toward automating anything done more than once, and discomfort with manual, undocumented changes

  • Experience with Windows Server administration, IIS, Active Directory/Entra ID, and Windows-based application stacks, which are critical for the on-prem component at the core of our product

  • Solid understanding of Azure networking (VNets, peering, private endpoints, NSGs), identity (Entra ID, RBAC, service principals and managed identities), and cost management tooling

  • Working knowledge of AWS core services (IAM, VPC, EC2/ECS, etc.) sufficient to support and extend an already well-architected environment

  • Experience integrating on-premises infrastructure with cloud environments (VPN/ExpressRoute, hybrid identity, hybrid networking)

  • A track record of introducing structure and standards into environments that grew organically or quickly; you enjoy imposing order, not just maintaining it

  • Strong scripting ability (PowerShell and/or Bash/Python)

  • Excellent judgment around security and access control; you think in terms of least privilege by default

  • Comfort operating with a high degree of autonomy and ownership, including making architectural calls and driving them through to adoption

**

Nice to Have: **

  • Familiarity with CI/CD tooling (Azure DevOps) for infrastructure pipelines

  • Prior experience leading a cloud environment from an ungoverned state to a well-architected one

  • Relevant certifications (Azure Solutions Architect, Azure Administrator, HashiCorp Terraform Associate)

**

What You'll Love About Dental Intelligence: **

  • Flexible Paid Time Off + 10 company-wide paid holidays

  • Competitive Medical, Dental & vision offerings, with buy up plan options, AND we match your HSA contributions.

  • Fully Paid Parental Leave

  • 401K Retirement savings plan with company match up to 5.5% of your earnings+ unlimited access to personal financial advisors.

  • Learning & Development Reimbursement program

  • Company paid Life, Disability & AD&D

  • Mental Health support programs, Cellphone & Gym membership Discounts, Corporate Sundance Passes, and more!

**

Please Note: ** All offers of employment are contingent upon successful completion of a background check, which may include verification of education, employment history, and other credentials. By applying, you confirm that all information you have provided is accurate and complete to the best of your knowledge, and you understand that misrepresentation may result in disqualification or termination.

** About Dental Intelligence **

We are the leading player in the SaaS analytics and workflow space for dental practices, launched in 2015 to help dentists manage and grow their practices. Our best-in-class tech makes it more fulfilling to be a dental professional and easier to be a patient. Nearly 9,000 dental practices utilize our platform to practice smarter, generating an average top-line production increase of 50% in the first 12 months. Whether a practice wants a comprehensive 2-year growth plan or simply a more effective Morning Huddle, we take the busy work out of growth. Our platform helps practices find patients, schedule them, follow up, collect payments, file their forms, design their treatment plans, and so much more.

We seek an experienced Sr Site Reliability Engineer who can help our organization to mature and scale the infrastructure behind our multi-cloud SaaS platform. If the profile below sounds like you - letโ€™s talk!