Haystack
← Back to Jobs
Technology
TS

Site Reliability Engineer (SRE)

TSQ Systems IncPhiladelphia, PA🇺🇸United StatesPosted Sep 30, 2026

Quick Overview

Seniority
Mid Senior
Work mode
Hybrid
Location
Philadelphia, PA, United States
Posted
18 hours ago
DockerShellAWSAnsibleAzureDatadogGoogle CloudGrafanaKubernetesPrometheusPythonTerraform

Job Description

Site Reliability Engineer (SRE)

Introduction:

The Site Reliability Engineer (SRE) will play a crucial role in ensuring the reliability and performance of our systems. They will work closely with the development and operations teams to implement monitoring and observability tools, automate processes, and provide support for high-availability environments.

Responsibilities:

  • Provide production support, reliability, and system performance expertise.
  • Utilize monitoring and observability tools such as Dynatrace, Datadog, Prometheus, or Grafana.
  • Work with cloud platforms (AWS/Google Cloud Platform/Azure) and containerization technologies (Docker, Kubernetes).
  • Handle incident management, root cause analysis (RCA), and participate in on-call support.
  • Automate tasks and processes using scripting languages like Python, Shell, or similar.
  • Implement CI/CD pipelines, DevOps practices, and infrastructure as code (Terraform/Ansible).

Requirements:

Required Skills:

  • 7+ years of experience in production support, reliability, and system performance.
  • Expertise in monitoring and observability tools like Dynatrace, Grafana.
  • Hands-on experience with cloud platforms (AWS/Google Cloud Platform/Azure) and containerization (Docker, Kubernetes).
  • Knowledge of incident management, root cause analysis (RCA), and on-call support in high-availability environments.
  • Experience in automation and scripting using Python, Shell, or similar for operational efficiency.
  • Familiarity with CI/CD pipelines, DevOps practices, and infrastructure as code (Terraform/Ansible).

Similar jobs