Role OverviewDesign, build, and mature the observability ecosystem that underpins FanDuel's platform and services. Partner with engineering and product teams to deliver scalable observability capabilities and serve as a subject matter expert in monitoring, alerting, and incident management.
What You Will Do
Deliver deep visibility into system behavior, combine system telemetry with user signals, and provide a holistic view of performance, reliability, and user experience. Explore how AI and machine learning can enhance observability and drive improvements in reliability, performance, and data-informed decision-making.
Why It Might Be a Fit
Contribute to the observability strategy and roadmap, design scalable observability solutions, and help establish best practices for monitoring, alerting, and incident management. Collaborate on cross-team initiatives to improve system reliability and apply automation and AI-assisted workflows to improve root cause analysis.
Requirements
- Solid hands-on experience in observability engineering, SRE, platform engineering, or related roles
- Strong expertise in monitoring and observability practices
- Experience contributing to observability or reliability initiatives across teams or services
- Proficiency with Kubernetes, cloud infrastructure (e.g. AWS), and infrastructure-as-code tools such as Terraform
- Ability to influence technical decisions within and across teams
- Good understanding of distributed systems principles and practical trade-offs
- Experience defining and implementing SLOs, SLIs, and alerting strategies
- Strong software engineering fundamentals and proficiency in at least one modern programming language
- Experience improving systems through automation and reducing operational toil
Benefits
- Medical, vision, and dental insurance
- Life insurance
- Disability insurance
- 401(k) matching program
- Paid personal time off
- 14 paid company holidays
- Paid sick time
- Annual bonus and long-term incentive opportunities
- Commuter benefits
- Pet insurance
]]>