Systems Engineer, AWS Incident Response

Amazon
Seattle, WA
Category Engineering
Job Description
Role Overview

As a Systems Engineer on AWS Incident Response, you will lead high-severity calls, triage complex failures, and drive incidents to mitigation in real time. You will also obsess over metrics and detection analysis, building dashboards and mechanisms that surface problems before customers notice.

What You Will Do

You will lead high-severity incident response calls, own operational processes, and use data to find the next improvement in how we detect and mitigate faster. You will also participate in an on-call rotation covering weekdays, weekends, and holidays.

Why It Might Be a Fit

This role requires experience in written and verbal communication skills, scripting, and troubleshooting complex technical problems. You will also have the opportunity to grow your development skills by taking on coding projects that accelerate incident response and reduce toil.

Requirements

  • 3+ years of systems engineering, or 3+ years of technical support experience
  • Experience in written and verbal communication skills to communicate with technical and non-technical audiences
  • Experience scripting in one or more language (e.g. Bash, Python, Perl, Ruby)
  • Understanding of operating systems (Linux), networking fundamentals, and distributed systems
  • Experience with operational monitoring, alerting, and metrics (CloudWatch, Datadog, Grafana, or equivalent)
  • Demonstrated ability to troubleshoot complex technical problems spanning multiple systems or services

Benefits

  • Dental insurance
  • Vision insurance
  • Prescription insurance
  • Basic Life & AD&D insurance
  • 401(k) matching
  • Paid time off
  • Parental leave
  • Sign-on payments
  • Restricted stock units (RSUs)
]]>