
Customer Reliability Engineer (Infrastructure)
Astronomer
Job description
-
The Astronomer Customer Reliability Engineering (CRE) team is responsible for the success of our customers’ usage of our managed Airflow service
-
The CREs are responsible for operating, monitoring, and maintaining the platform to ensure availability, predictability, and reliable operations
-
As an infrastructure specialist within the team, you will focus on the reliability of the underlying cloud infrastructure and Kubernetes clusters
-
This entails responding to incidents either raised by a customer, or from our monitoring system and then taking further steps to ensure problems are permanently resolved or monitored
-
As owners of the observability platform, CRE has unlimited potential to improve the reliability of the product and deliver the best possible outcome for our customers
-
This role is directly customer-facing and gives exposure to very diverse problems and requirements
-
CRE get the opportunity to interface with customers from a variety of industries across different cloud providers, and all with different expectations
-
Your contributions will directly impact customers’ success with using the Astronomer products, and you will be able to help make meaningful improvements to the customer experience
-
Provide solutions to customers to make them successful using our products
-
Troubleshoot customer environments and engage in active triaging with customers
-
Participate in on-call rotation for weekend coverage
-
Provide feedback to the product development teams on customer needs and pain points
-
Build out our monitoring and alerting systems
-
Build and maintain automation to ensure daily operational tasks are handled as efficiently as possible
-
Help direct the architecture of the products and contribute where possible
-
Own the customer experience, working directly with customers to prioritize and solve issues, meet SLAs, and provide “white glove” guidance on the path to production
-
Participate remotely within a fully distributed team
-
Enhance and enrich customer documentation
-
Work with the latest technology and multi-cloud implementations
Benefits
-
Remote friendly, work from anywhere
-
If you’re in a city with an office, it’s there (and stocked with snacks) when you need it
-
Co-workers in over 40 states and 15 countries around the world
-
Health, dental, and vision insurance at little to no cost for individuals, and at competitive rates for your dependents
-
Disability and life insurance policies in case something happens
-
Unlimited vacation - we do track vacation days and actively encourage people to take it (on average, Astronomers take 20 days of vacation per year)
-
Parental leave
-
Laptop of your choice, and a stipend to help with your work-from-home setup
-
Monthly $170 pre-tax stipend intended to cover your cellphone and WiFi bills
-
Yearly internal summit for all employees
-
Regular access to conferences, workshops, and meetups in the ecosystem- 5 years of experience, preferably with large, complex cloud infrastructures operating at scale
-
DevOps or CI/CD experience
-
Experience managing a Production distributed system with at least one major cloud provider (one or all: AWS, GCP, Azure)
-
Python scripting
-
Strong Linux experience
-
Strong communication skills
-
Knowledge of how to operate and monitor issues for distributed systems
-
3 years of experience with Kubernetes
-
Previous experience in handling customers issues (internal or external)
-
Good troubleshooting Skills
-
Worked with Kubernetes Custom Resources
-
Depth of knowledge with Azure
-
Airflow/Big Data Orchestration experience
-
IaC experience
-
Experience as a Site Reliability Engineer