Role OverviewAgileEngine is an Inc. 5000 company that creates award-winning software for Fortune 500 brands and trailblazing startups. We rank among the leaders in areas like application development and AI/ML, and our people-first culture has earned us multiple Best Place to Work awards. We are looking for a DevOps / Site Reliability Engineer to maintain operational resilience and 24/7 stability for a multi-cloud enterprise security program.
What You Will Do
Scale and maintain the ability to drive operational stability across multi-cloud environments, engineer unified security policies, design and maintain CI/CD pipelines, act on continuous monitoring alerts, and serve as Incident Commander on major and critical incidents.
Why It Might Be a Fit
If you're looking for a place to grow, make an impact, and work with people who care, we'd love to meet you! You will drive major-incident calls, own post-incident remediation follow-through, draft stakeholder communications, and develop divisional incident-management playbooks alongside multi-cloud security guardrails.
Requirements
- 5+ years of SRE experience
- Hands-on incident command in a 24/7 financial services environment
- In-depth architectural expertise in multi-cloud defense, federated IAM, and zero-trust principles
- Strong practical experience with Kubernetes, Terraform, CI/CD orchestration, and Python/Go scripting
- Senior-level, hands-on incident-command experience driving major/critical incident calls to resolution in a 24x7 production environment
- Proven track record of remediation follow-up
- Demonstrated skill drafting and issuing incident notification communications to both technical and executive audiences
- Direct experience authoring divisional/group incident-management playbooks and escalation procedures
Benefits
- Growth without limits
- Competitive compensation
- Flexibility
- Meaningful, modern projects
- Collaborative culture
- Well-being & support
- Access to local well-being programs
- People-focused support tailored to your location
]]>