DevOps / Site Reliability Engineer

Agile Engine
Any Location, TX
Job Description
Role Overview

We are looking for a DevOps / Site Reliability Engineer to maintain operational resilience and 24/7 stability for a multi-cloud enterprise security program. The role requires 5+ years of SRE experience with hands-on incident command in a 24/7 financial services environment.

What You Will Do

Scale and maintain the ability to drive operational stability across multi-cloud environments, engineer unified security policies, design and maintain CI/CD pipelines, and serve as Incident Commander on major and critical incidents.

Why It Might Be a Fit

You will drive major-incident calls, own post-incident remediation follow-through, draft stakeholder communications, and develop divisional incident-management playbooks alongside multi-cloud security guardrails.

Requirements

  • Authorized to work for ANY employer in the US
  • 5+ years of experience
  • In-depth architectural expertise in multi-cloud defense, federated IAM, and zero-trust principles
  • Strong practical experience with Kubernetes, Terraform, CI/CD orchestration, and Python/Go scripting
  • Senior-level, hands-on incident-command experience driving major/critical incident calls to resolution in a 24x7 production environment
  • Proven track record of remediation follow-up
  • Demonstrated skill drafting and issuing incident notification communications to both technical and executive audiences
  • Direct experience authoring divisional/group incident-management playbooks and escalation procedures
  • Fully autonomous
  • Extensive experience deploying and tuning APIs from modern CNAPP/CSPM platforms, ideally Wiz
  • Prior experience building platforms subject to strict financial compliance standards (PCI-DSS, SOC2)
  • Upper-intermediate English level

Benefits

  • Growth without limits
  • Competitive compensation
  • Flexibility
  • Meaningful, modern projects
  • Collaborative culture
  • Well-being & support
]]>