Role OverviewStack AV Site Reliability Engineers are responsible for enabling and ensuring our production systems meet their service-level objectives. Through the implementation of centralized observability and automation, the SRE team constantly ensures the health, reliability, scalability, and performance of Stack AV’s infrastructure.
What You Will Do
Instrument systems scheduling and executing large-scale batch workloads across Kubernetes clusters. Diagnose and triage job failures for customers. Collaborate with teams across the company to understand workload requirements and improve platform capabilities.
Why It Might Be a Fit
Strong experience with Kubernetes and container orchestration in production grade environments. Strong communication skills and the ability to work effectively in a diverse and distributed team.
Requirements
- Fundamental understanding of Linux operating system internals, TCP/IP networking, and storage subsystems.
- Strong experience implementing and debugging cloud native and open source tools such as Kubernetes, etcd, Prometheus, OpenTelemetry.
- Strong communication skills and the ability to work effectively in a diverse and distributed team.
Benefits
- equal opportunity workplace
- diverse teams produce the best ideas and outcomes
- culture of inclusion, entrepreneurship, and innovation
]]>