← Back to Jobs
Technology
Lead Site Reliability Engineer (F2F - NY)
Keen Technology Solutions LLCBuffalo, NY🇺🇸United StatesPosted 7 Aug 2026
Quick Overview
Work Type
On Site
Level
Mid Senior
Job Description
Job Title : Lead Site Reliability Engineer
Location: Buffalo, NY 100% onsite
Duration: Long Term
Primary Responsibilities
- Design, implement, and support highly available, scalable, and resilient applications and cloud infrastructure following enterprise technology standards and SRE best practices.
- Lead initiatives to improve system reliability, availability, performance, and operational maturity through automation and engineering excellence.
- Define, implement, and monitor Service Level Objectives (SLOs), Service Level Indicators (SLIs), and error budgets for critical business services.
- Develop comprehensive observability strategies leveraging Dynatrace, OpenTelemetry (OTel), distributed tracing, metrics, logging, dashboards, and alerting solutions.
- Design and maintain end-to-end monitoring solutions that provide actionable insights into application, infrastructure, and customer experience health.
- Analyze production telemetry to proactively identify performance bottlenecks, reliability risks, and capacity constraints.
- Lead incident response activities for high-severity production events, coordinating cross-functional teams to restore services and minimize customer impact.
- Perform and facilitate Root Cause Analysis (RCA) activities, ensuring corrective and preventive actions are identified, prioritized, and implemented.
- Drive operational excellence through automation of repetitive tasks, operational workflows, deployments, recovery procedures, and reliability controls.
- Partner with development teams to build reliable and observable services throughout the Software Development Lifecycle (SDLC).
- Design, develop, and execute automated regression testing strategies to validate application stability, reliability, and performance following deployments and infrastructure changes.
- Review test coverage and reliability validation approaches to ensure comprehensive testing and risk mitigation.
- Create, maintain, and improve Infrastructure as Code (IaC) solutions using Terraform for cloud infrastructure provisioning, configuration management, and environment standardization.
- Support and optimize Microsoft Azure environments, including Azure App Services, resource management, scaling strategies, deployment automation, and application lifecycle management.
- Utilize Azure-native tools such as Azure Monitor, Application Insights, Log Analytics, and related services to improve platform visibility and reliability.
- Drive implementation of performance testing, resiliency testing, fault tolerance validation, and disaster recovery preparedness within assigned domains.
- Establish operational readiness standards and ensure applications meet reliability, scalability, observability, and supportability requirements before production deployment.
- Review architectural designs and provide recommendations to improve platform resiliency, operational efficiency, and cloud optimization.
- Lead capacity planning, performance tuning, and workload optimization efforts across production environments.
- Develop and maintain operational runbooks, incident playbooks, knowledge articles, and standard operating procedures.
- Serve as a key partner with engineering, infrastructure, cybersecurity, architecture, and support teams to identify and implement continuous process improvements spanning organizational boundaries.
- Communicate system health, reliability trends, operational risks, and remediation strategies to technical and business stakeholders.
- Present reliability initiatives, operational metrics, and engineering recommendations at architecture reviews, technical forums, and leadership meetings.
- Mentor engineers on observability, cloud engineering, automation, SRE principles, and operational best practices.
- Understand and adhere to the Company's risk and regulatory standards, policies, and controls in accordance with the Company's Risk Appetite.
- Identify reliability, operational, and technology risks requiring escalation to management.
- Promote an environment that supports a culture of belonging and reflects the M&T Bank brand.
- Maintain M&T internal control standards, including timely implementation of internal and external audit findings and regulatory requirements as applicable.
- Complete other related duties as assigned.
Core Requirements
- Strong experience in observability and monitoring, including hands-on expertise with:
- Dynatrace
- OpenTelemetry (OTel)
- Distributed tracing
- Metrics collection and analysis
- Centralized logging and log aggregation
- Alerting and dashboard development
- Proven experience designing and executing automated regression testing frameworks and test suites to ensure application and platform stability following deployments.
- Strong proficiency in Infrastructure as Code (IaC) using Terraform.
- Experience with CI/CD pipelines, deployment automation, and operational tooling.
- Expert knowledge of production systems monitoring, incident management, and operational troubleshooting.
- Strong understanding of application performance management, distributed systems, and modern cloud-native architectures.
Cloud & Platform Expertise
- Strong experience with Microsoft Azure, including:
- Azure App Services
- Resource Groups
- Azure networking concepts
- Scaling and performance optimization
- Deployment and release management
- Application lifecycle management
- Experience leveraging Azure-native operational tooling such as:
- Azure Monitor
- Application Insights
- Log Analytics
- Azure dashboards and alerting
- Experience supporting cloud-native and hybrid infrastructure environments.
Reliability & Engineering Practices
- Demonstrated experience implementing and operating SRE practices, including:
- Service Level Objectives (SLOs)
- Service Level Indicators (SLIs)
- Error budgets
- Incident management
- Problem management
- Root Cause Analysis (RCA)
- Reliability automation
- Ability to improve system reliability through:
- Performance tuning
- Capacity planning
- Observability-driven insights
- Proactive issue detection
- Reliability engineering initiatives
- Experience developing automated recovery mechanisms and self-healing solutions.
- Knowledge of resiliency engineering patterns, disaster recovery planning, and high-availability architectures.
Skills
Azure
Terraform
Similar jobs
Site Reliability Engineer
Stellent IT LLC · New York, United States
9 minutes agoPlatform Engineer III
Apex Systems · Cincinnati, United States
32 minutes ago$60 - $68/hrDevOps Engineer - Madison, WI, Columbus, OH, Chicago, IL, Minneapolis, MN, Detroit, MI.
TechniPros, LLC · Madison, United States
34 minutes agoDevOps Engineer - Dallas, TX, Austin, TX, Houston, TX, San Antonio, TX.
TechniPros, LLC · Dallas, United States
36 minutes agoSalesforce DevOps Engineer
Tixy Services LLC · Dallas, United States
55 minutes agoDevOps Engineer (Java Environment)
Wise Skulls Corp. · Austin, United States
55 minutes ago