Astronomer logo

Customer Reliability Engineer (Infrastructure)

Astronomer

Remotesenior$125k–$130kPosted 11h ago

Job description

  • The Astronomer Customer Reliability Engineering (CRE) team is responsible for the success of our customers’ usage of our managed Airflow service

  • The CREs are responsible for operating, monitoring, and maintaining the platform to ensure availability, predictability, and reliable operations

  • As an infrastructure specialist within the team, you will focus on the reliability of the underlying cloud infrastructure and Kubernetes clusters

  • This entails responding to incidents either raised by a customer, or from our monitoring system and then taking further steps to ensure problems are permanently resolved or monitored

  • As owners of the observability platform, CRE has unlimited potential to improve the reliability of the product and deliver the best possible outcome for our customers

  • This role is directly customer-facing and gives exposure to very diverse problems and requirements

  • CRE get the opportunity to interface with customers from a variety of industries across different cloud providers, and all with different expectations

  • Your contributions will directly impact customers’ success with using the Astronomer products, and you will be able to help make meaningful improvements to the customer experience

  • Provide solutions to customers to make them successful using our products

  • Troubleshoot customer environments and engage in active triaging with customers

  • Participate in on-call rotation for weekend coverage

  • Provide feedback to the product development teams on customer needs and pain points

  • Build out our monitoring and alerting systems

  • Build and maintain automation to ensure daily operational tasks are handled as efficiently as possible

  • Help direct the architecture of the products and contribute where possible

  • Own the customer experience, working directly with customers to prioritize and solve issues, meet SLAs, and provide “white glove” guidance on the path to production

  • Participate remotely within a fully distributed team

  • Enhance and enrich customer documentation

  • Work with the latest technology and multi-cloud implementations

Benefits

  • Remote friendly, work from anywhere

  • If you’re in a city with an office, it’s there (and stocked with snacks) when you need it

  • Co-workers in over 40 states and 15 countries around the world

  • Health, dental, and vision insurance at little to no cost for individuals, and at competitive rates for your dependents

  • Disability and life insurance policies in case something happens

  • Unlimited vacation - we do track vacation days and actively encourage people to take it (on average, Astronomers take 20 days of vacation per year)

  • Parental leave

  • Laptop of your choice, and a stipend to help with your work-from-home setup

  • Monthly $170 pre-tax stipend intended to cover your cellphone and WiFi bills

  • Yearly internal summit for all employees

  • Regular access to conferences, workshops, and meetups in the ecosystem- 5 years of experience, preferably with large, complex cloud infrastructures operating at scale

  • DevOps or CI/CD experience

  • Experience managing a Production distributed system with at least one major cloud provider (one or all: AWS, GCP, Azure)

  • Python scripting

  • Strong Linux experience

  • Strong communication skills

  • Knowledge of how to operate and monitor issues for distributed systems

  • 3 years of experience with Kubernetes

  • Previous experience in handling customers issues (internal or external)

  • Good troubleshooting Skills

  • Worked with Kubernetes Custom Resources

  • Depth of knowledge with Azure

  • Airflow/Big Data Orchestration experience

  • IaC experience

  • Experience as a Site Reliability Engineer