Domino Data Lab logo

Staff Platform Reliability Engineer

Domino Data Lab

Remotelead$185k–$230kPosted 8h ago

Job description

  • The Automation Team at Domino acts as a force multiplier for engineering, building the tools and systems that enable teams to ship code confidently and consistently

  • Focused on iteration and continuous improvement, the team looks for targeted enhancements that can create outsized impact across the organization

  • They are responsible for developing and maintaining the frameworks and tooling that empower developers and Quality Engineers to create meaningful, actionable end-to-end tests for their features

  • By leveraging Domino’s own platform to build powerful visualizations and streamline workflows, the team helps accelerate issue identification and resolution, ultimately improving product quality and delivery speed

  • In your first year, you will:

  • Serve as the technical owner of Tempest, Domino’s scale testing platform tools, ensuring they remain reliable, extensible, and aligned with evolving product needs

  • Evolve and modernize scale testing infrastructure to support continued growth and increasing product complexity

  • Deliver accurate, data-driven sizing recommendations for customer-facing documentation based on rigorous empirical testing

  • Dramatically improve validation and reporting within scale testing by introducing repeatable, automated validation criteria that reduce manual diagnostics while increasing confidence in results

  • Establish and operationalize scale testing on cloud platforms, ensuring appropriate sizing and configuration guidance for this increasingly divergent product line

  • Partner with platform teams to enable effective scale testing across additional cloud providers, helping position Domino for future multi-cloud success

  • Increase the efficiency and leverage of a small team by building automation that scales operationally as the product and customer base grow

Benefits

  • Health & wellness: Premium medical, dental, and vision insurance plans, including a free option for you and your family

  • Commuter benefits: What’s your transportation of choice? We support your daily commute, whether Uber, Lyft, bus, or train

  • Love for parents: New parents receive up to eight weeks of fully paid-parental leave. P.S. Got baby pics? Share them in our Slack channel dedicated to kids

  • Flexible paid time off: We value your right to disconnect. Take the time you need to recharge, spend time with family, and explore the world

  • Annual education reimbursement: Pursue the professional growth that will help you excel in your role with educational opportunities that work for your needs, and at your speed

  • Own your outcome: Feel supported to get your work done how and when you need to, regardless of your location- Strong proficiency in Python and comfort working in a large, modular codebase that spans orchestration, infrastructure automation, and systems integration

  • Familiarity with performance and load testing methodologies (e.g., Locust, k6, or similar) as part of a broader infrastructure or reliability practice

  • Demonstrated ability to go beyond detection to resolution: profiling services, identifying resource bottlenecks, and working with engineering teams to ship durable fixes

  • Experience with observability stacks (Prometheus, Grafana, New Relic, or similar) — writing queries, building dashboards, and using metrics to diagnose performance and reliability issues at the systems level

  • Background in SRE, platform engineering, or infrastructure with hands-on experience operating and troubleshooting distributed systems in production Kubernetes environments

  • Clear ownership mindset — self-directed, accountable, and able to communicate priorities and status effectively in a remote, async environment