Lead Site Reliability Engineer

Glint Tech Solutions
Buffalo, NY
Job Description
Role Overview

The Lead Site Reliability Engineer is responsible for ensuring the reliability, scalability, performance, and operational excellence of critical banking platforms and applications. This senior individual contributor will design, implement, and improve SRE practices across the software development lifecycle, working closely with application development, infrastructure, platform engineering, and business teams to enhance system resiliency through automation, observability, testing, and proactive operational management.

What You Will Do

Design, implement, and support highly available, scalable, and resilient applications and cloud infrastructure following enterprise SRE best practices. Define, implement, and monitor SLOs, SLIs, and error budgets for critical business services. Develop observability strategies using Dynatrace, OpenTelemetry (OTel), distributed tracing, metrics, logging, dashboards, and alerting.

Why It Might Be a Fit

The ideal candidate will have strong hands-on experience with Dynatrace, OpenTelemetry (OTel), distributed tracing, metrics collection, and centralized logging. Proven experience designing and executing automated regression testing frameworks is also required. The candidate should have a strong proficiency in Infrastructure as Code (IaC) using Terraform and experience with CI/CD pipelines, deployment automation, and operational tooling.

Requirements

  • Strong hands-on experience with Dynatrace, OpenTelemetry (OTel), distributed tracing, metrics collection, and centralized logging
  • Proven experience designing and executing automated regression testing frameworks
  • Strong proficiency in Infrastructure as Code (IaC) using Terraform
  • Experience with CI/CD pipelines, deployment automation, and operational tooling
  • Expert knowledge of production systems monitoring, incident management, and operational troubleshooting
  • Strong understanding of application performance management, distributed systems, and cloud-native architectures
  • Strong experience with Microsoft Azure (App Services, Resource Groups, networking, scaling, deployment/release management)
  • Experience with Azure Monitor, Application Insights, Log Analytics, and Azure dashboards/alerting
  • Experience supporting cloud-native and hybrid infrastructure environments
  • Demonstrated experience implementing SRE practices — SLOs, SLIs, error budgets, incident/problem management, RCA, reliability automation
  • Ability to improve system reliability through performance tuning, capacity planning, and observability-driven insights
  • Experience developing automated recovery mechanisms and self-healing solutions
  • Knowledge of resiliency engineering patterns, disaster recovery planning, and high-availability architectures

Benefits

  • Travel expenses covered by the company
  • Paid holidays
  • Equity
  • Healthcare
  • PTO
  • Retirement
  • Learning budget
  • Parental leave
  • Wellness
  • Visa/relocation assistance
  • Remote flexibility
  • Stipends
  • Bonus/commission
]]>