Role OverviewThe Astronomer Customer Reliability Engineering (CRE) team is responsible for the success of our customers' usage of our managed Airflow service. As an infrastructure specialist within the team, you will focus on the reliability of the underlying cloud infrastructure and Kubernetes clusters.
What You Will Do
Provide solutions to customers to make them successful using our products, troubleshoot customer environments, and engage in active triaging with customers. Participate in on-call rotation for weekend coverage, provide feedback to the product development teams on customer needs and pain points.
Why It Might Be a Fit
This role is directly customer-facing and gives exposure to very diverse problems and requirements. CRE get the opportunity to interface with customers from a variety of industries across different cloud providers, and all with different expectations.
Requirements
- 5 years of experience, preferably with large, complex cloud infrastructures operating at scale
- 3 years of experience with Kubernetes
- Experience managing a Production distributed system with at least one major cloud provider (one or all: AWS, GCP, Azure)
- Strong Linux experience
- Knowledge of how to operate and monitor issues for distributed systems
- Previous experience in handling customers issues (internal or external)
- Strong communication skills
- DevOps or CI/CD experience
- Python scripting
- Good troubleshooting Skills
Benefits
- Comprehensive benefits package
- Equity component
]]>