
Staff Platform Reliability Engineer
Domino Data Lab
Job description
-
The Automation Team at Domino acts as a force multiplier for engineering, building the tools and systems that enable teams to ship code confidently and consistently
-
Focused on iteration and continuous improvement, the team looks for targeted enhancements that can create outsized impact across the organization
-
They are responsible for developing and maintaining the frameworks and tooling that empower developers and Quality Engineers to create meaningful, actionable end-to-end tests for their features
-
By leveraging Domino’s own platform to build powerful visualizations and streamline workflows, the team helps accelerate issue identification and resolution, ultimately improving product quality and delivery speed
-
In your first year, you will:
-
Serve as the technical owner of Tempest, Domino’s scale testing platform tools, ensuring they remain reliable, extensible, and aligned with evolving product needs
-
Evolve and modernize scale testing infrastructure to support continued growth and increasing product complexity
-
Deliver accurate, data-driven sizing recommendations for customer-facing documentation based on rigorous empirical testing
-
Dramatically improve validation and reporting within scale testing by introducing repeatable, automated validation criteria that reduce manual diagnostics while increasing confidence in results
-
Establish and operationalize scale testing on cloud platforms, ensuring appropriate sizing and configuration guidance for this increasingly divergent product line
-
Partner with platform teams to enable effective scale testing across additional cloud providers, helping position Domino for future multi-cloud success
-
Increase the efficiency and leverage of a small team by building automation that scales operationally as the product and customer base grow
Benefits
-
Health & wellness: Premium medical, dental, and vision insurance plans, including a free option for you and your family
-
Commuter benefits: What’s your transportation of choice? We support your daily commute, whether Uber, Lyft, bus, or train
-
Love for parents: New parents receive up to eight weeks of fully paid-parental leave. P.S. Got baby pics? Share them in our Slack channel dedicated to kids
-
Flexible paid time off: We value your right to disconnect. Take the time you need to recharge, spend time with family, and explore the world
-
Annual education reimbursement: Pursue the professional growth that will help you excel in your role with educational opportunities that work for your needs, and at your speed
-
Own your outcome: Feel supported to get your work done how and when you need to, regardless of your location- Strong proficiency in Python and comfort working in a large, modular codebase that spans orchestration, infrastructure automation, and systems integration
-
Familiarity with performance and load testing methodologies (e.g., Locust, k6, or similar) as part of a broader infrastructure or reliability practice
-
Demonstrated ability to go beyond detection to resolution: profiling services, identifying resource bottlenecks, and working with engineering teams to ship durable fixes
-
Experience with observability stacks (Prometheus, Grafana, New Relic, or similar) — writing queries, building dashboards, and using metrics to diagnose performance and reliability issues at the systems level
-
Background in SRE, platform engineering, or infrastructure with hands-on experience operating and troubleshooting distributed systems in production Kubernetes environments
-
Clear ownership mindset — self-directed, accountable, and able to communicate priorities and status effectively in a remote, async environment