Haystack
← Back to Jobs
Technology

Lead Site Reliability Engineer (F2F - NY)

Keen Technology Solutions LLCBuffalo, NY🇺🇸United StatesPosted 7 Aug 2026

Quick Overview

Work Type
On Site
Level
Mid Senior

Job Description

Job Title : Lead Site Reliability Engineer

Location: Buffalo, NY 100% onsite

Duration: Long Term

Primary Responsibilities

  • Design, implement, and support highly available, scalable, and resilient applications and cloud infrastructure following enterprise technology standards and SRE best practices.
  • Lead initiatives to improve system reliability, availability, performance, and operational maturity through automation and engineering excellence.
  • Define, implement, and monitor Service Level Objectives (SLOs), Service Level Indicators (SLIs), and error budgets for critical business services.
  • Develop comprehensive observability strategies leveraging Dynatrace, OpenTelemetry (OTel), distributed tracing, metrics, logging, dashboards, and alerting solutions.
  • Design and maintain end-to-end monitoring solutions that provide actionable insights into application, infrastructure, and customer experience health.
  • Analyze production telemetry to proactively identify performance bottlenecks, reliability risks, and capacity constraints.
  • Lead incident response activities for high-severity production events, coordinating cross-functional teams to restore services and minimize customer impact.
  • Perform and facilitate Root Cause Analysis (RCA) activities, ensuring corrective and preventive actions are identified, prioritized, and implemented.
  • Drive operational excellence through automation of repetitive tasks, operational workflows, deployments, recovery procedures, and reliability controls.
  • Partner with development teams to build reliable and observable services throughout the Software Development Lifecycle (SDLC).
  • Design, develop, and execute automated regression testing strategies to validate application stability, reliability, and performance following deployments and infrastructure changes.
  • Review test coverage and reliability validation approaches to ensure comprehensive testing and risk mitigation.
  • Create, maintain, and improve Infrastructure as Code (IaC) solutions using Terraform for cloud infrastructure provisioning, configuration management, and environment standardization.
  • Support and optimize Microsoft Azure environments, including Azure App Services, resource management, scaling strategies, deployment automation, and application lifecycle management.
  • Utilize Azure-native tools such as Azure Monitor, Application Insights, Log Analytics, and related services to improve platform visibility and reliability.
  • Drive implementation of performance testing, resiliency testing, fault tolerance validation, and disaster recovery preparedness within assigned domains.
  • Establish operational readiness standards and ensure applications meet reliability, scalability, observability, and supportability requirements before production deployment.
  • Review architectural designs and provide recommendations to improve platform resiliency, operational efficiency, and cloud optimization.
  • Lead capacity planning, performance tuning, and workload optimization efforts across production environments.
  • Develop and maintain operational runbooks, incident playbooks, knowledge articles, and standard operating procedures.
  • Serve as a key partner with engineering, infrastructure, cybersecurity, architecture, and support teams to identify and implement continuous process improvements spanning organizational boundaries.
  • Communicate system health, reliability trends, operational risks, and remediation strategies to technical and business stakeholders.
  • Present reliability initiatives, operational metrics, and engineering recommendations at architecture reviews, technical forums, and leadership meetings.
  • Mentor engineers on observability, cloud engineering, automation, SRE principles, and operational best practices.
  • Understand and adhere to the Company's risk and regulatory standards, policies, and controls in accordance with the Company's Risk Appetite.
  • Identify reliability, operational, and technology risks requiring escalation to management.
  • Promote an environment that supports a culture of belonging and reflects the M&T Bank brand.
  • Maintain M&T internal control standards, including timely implementation of internal and external audit findings and regulatory requirements as applicable.
  • Complete other related duties as assigned.

Core Requirements

  • Strong experience in observability and monitoring, including hands-on expertise with:
  • Dynatrace
  • OpenTelemetry (OTel)
  • Distributed tracing
  • Metrics collection and analysis
  • Centralized logging and log aggregation
  • Alerting and dashboard development
  • Proven experience designing and executing automated regression testing frameworks and test suites to ensure application and platform stability following deployments.
  • Strong proficiency in Infrastructure as Code (IaC) using Terraform.
  • Experience with CI/CD pipelines, deployment automation, and operational tooling.
  • Expert knowledge of production systems monitoring, incident management, and operational troubleshooting.
  • Strong understanding of application performance management, distributed systems, and modern cloud-native architectures.

Cloud & Platform Expertise

  • Strong experience with Microsoft Azure, including:
  • Azure App Services
  • Resource Groups
  • Azure networking concepts
  • Scaling and performance optimization
  • Deployment and release management
  • Application lifecycle management
  • Experience leveraging Azure-native operational tooling such as:
  • Azure Monitor
  • Application Insights
  • Log Analytics
  • Azure dashboards and alerting
  • Experience supporting cloud-native and hybrid infrastructure environments.

Reliability & Engineering Practices

  • Demonstrated experience implementing and operating SRE practices, including:
  • Service Level Objectives (SLOs)
  • Service Level Indicators (SLIs)
  • Error budgets
  • Incident management
  • Problem management
  • Root Cause Analysis (RCA)
  • Reliability automation
  • Ability to improve system reliability through:
  • Performance tuning
  • Capacity planning
  • Observability-driven insights
  • Proactive issue detection
  • Reliability engineering initiatives
  • Experience developing automated recovery mechanisms and self-healing solutions.
  • Knowledge of resiliency engineering patterns, disaster recovery planning, and high-availability architectures.

Skills

Azure
Terraform

Similar jobs