Role OverviewAs a Systems Engineer on AWS Incident Response, you will lead high-severity calls, triage complex failures, and drive incidents to mitigation in real time. You will also obsess over metrics and detection analysis, building dashboards and mechanisms that surface problems before customers notice.
What You Will Do
You will lead high-severity incident response calls, own operational processes, and use data to find the next improvement in how we detect and mitigate faster. You will also participate in an on-call rotation covering weekdays, weekends, and holidays.
Why It Might Be a Fit
This role requires experience in written and verbal communication skills, scripting, and troubleshooting complex technical problems. You will also have the opportunity to grow your development skills by taking on coding projects that accelerate incident response and reduce toil.
Requirements
- 3+ years of systems engineering, or 3+ years of technical support experience
- Experience in written and verbal communication skills to communicate with technical and non-technical audiences
- Experience scripting in one or more language (e.g. Bash, Python, Perl, Ruby)
- Understanding of operating systems (Linux), networking fundamentals, and distributed systems
- Experience with operational monitoring, alerting, and metrics (CloudWatch, Datadog, Grafana, or equivalent)
- Demonstrated ability to troubleshoot complex technical problems spanning multiple systems or services
Benefits
- Dental insurance
- Vision insurance
- Prescription insurance
- Basic Life & AD&D insurance
- 401(k) matching
- Paid time off
- Parental leave
- Sign-on payments
- Restricted stock units (RSUs)
]]>