Haystack
← Back to Jobs
Technology
SA

SRE

SLK America Inc.Buffalo, NY🇺🇸United StatesPosted Sep 22, 2026

Quick Overview

Seniority
Mid Senior
Work mode
On Site
Location
Buffalo, NY, United States
Posted
22 hours ago
DockerMicroservicesAWSELKLoad BalancingNew RelicSplunkAnsibleAzureBashCloudFormationDNSDatadogGitHub ActionsGitLab CIGoogle CloudGrafanaJenkinsKubernetesPowerShellPrometheusPythonTerraformVMware

Job Description

Site Reliability Engineer (SRE) - Job Description

Job Title

Site Reliability Engineer (SRE)

Job Summary

We are seeking a highly motivated and skilled Site Reliability Engineer (SRE) to join our engineering team. The ideal candidate will be responsible for maintaining the reliability, scalability, performance, and security of mission-critical applications and infrastructure. The SRE will work closely with development, cloud, and operations teams to automate processes, improve system availability, and drive operational excellence.

Key Responsibilities

  • Design, implement, and maintain highly available and scalable infrastructure.
  • Monitor application and infrastructure health using observability tools.
  • Automate operational tasks using scripting and Infrastructure as Code (IaC).
  • Manage incident response, root cause analysis, and post-incident reviews.
  • Define and monitor Service Level Indicators (SLIs), Service Level Objectives (SLOs), and Service Level Agreements (SLAs).
  • Optimize system performance, reliability, and resource utilization.
  • Build and maintain CI/CD pipelines for seamless application deployments.
  • Collaborate with software engineering teams to improve application resilience.
  • Implement disaster recovery, backup, and business continuity strategies.
  • Ensure security best practices across cloud and on-premises environments.
  • Create and maintain technical documentation, runbooks, and operational procedures.
  • Participate in on-call support and production issue management.

Required Qualifications

  • Bachelor's degree in Computer Science, Information Technology, Engineering, or related field.
  • 4+ years of experience in Site Reliability Engineering, DevOps, Cloud Engineering, or Systems Administration.
  • Strong experience with Linux/Unix systems administration.
  • Hands-on experience with cloud platforms such as AWS, Azure, or Google Cloud Platform (Google Cloud Platform).
  • Proficiency in scripting/programming languages such as Python, Bash, Go, or PowerShell.
  • Experience with containerization technologies such as Docker and Kubernetes.
  • Knowledge of CI/CD tools such as Jenkins, GitHub Actions, GitLab CI/CD, or Azure DevOps.
  • Experience with Infrastructure as Code tools such as Terraform, CloudFormation, or Ansible.
  • Strong understanding of networking, DNS, load balancing, and security principles.
  • Experience with monitoring tools such as Prometheus, Grafana, Datadog, New Relic, Splunk, or ELK Stack.

Preferred Qualifications

  • Experience supporting large-scale distributed systems.
  • Cloud certifications (AWS, Azure, or Google Cloud Platform).
  • Kubernetes Administrator (CKA) or similar certification.
  • Understanding of microservices architecture.
  • Experience with Chaos Engineering and Reliability Engineering practices.
  • Knowledge of database administration and performance tuning.

Technical Skills

Cloud & Infrastructure

  • AWS, Azure, Google Cloud Platform
  • VMware, Linux, Unix

Automation & IaC

  • Terraform
  • Ansible
  • CloudFormation
  • PowerShell
  • Bash
  • Python

Containers & Orchestration

  • Docker
  • Kubernetes
  • OpenShift

Monitoring & Observability

  • Prometheus
  • Grafana
  • Datadog
  • Splunk
  • ELK Stack
  • New Relic

CI/CD

  • Jenkins
  • Azure DevOps
  • GitHub Actions
  • GitLab CI/CD

Key Competencies

  • Problem-solving and troubleshooting
  • Incident management
  • Automation mindset
  • Strong communication and collaboration skills
  • Analytical thinking
  • Ownership and accountability
  • Continuous improvement mindset

Success Metrics

  • System uptime and availability
  • Mean Time to Detect (MTTD)
  • Mean Time to Resolve (MTTR)
  • Service reliability and performance
  • Deployment success rate
  • Reduction in operational toil through automation

Location: Remote/Hybrid/Onsite
Experience: 4-10+ Years
Employment Type: Full-Time

This JD can be used for hiring Mid-Level to Senior SRE professionals in enterprise cloud and production support environments.

Similar jobs