Haystack
← Back to Jobs
Technology
AG

Site Reliability Engineer (SRE)

ASCII Group LLCAustin, TX🇺🇸United StatesPosted Oct 8, 2026

Quick Overview

Seniority
Mid Senior
Work mode
On Site
Location
Austin, TX, United States
Posted
21 hours ago
AWSCloudFormationGoKubernetesPythonTerraform

Job Description

Hi ,

 

The following requirement is open with our client. 

Title                                       : Site Reliability Engineer (SRE)

Location                               : Austin, TX (Onsite)

Duration                              : 12 Months

Relevant Experience      : 8+

C2C or W2

Job Responsibilities:       

  • ·         We are looking for a skilled Site Reliability Engineer (SRE) to design, build, and maintain highly available, scalable, secure, and reliable cloud infrastructure and applications.
  • ·         Must have Apple Exp.
  • ·         Should be able to do coding independently in Python and GoLang, The ideal candidate will have strong hands-on experience with AWS, Kubernetes, Python, Linux, and cloud-native technologies.
  • ·         You will work closely with Development, DevOps, Security, and Operations teams to improve system reliability, automation, observability, and operational efficiency.
  • ·         Key Responsibilities
  • ·         Design, implement, and maintain highly available and scalable infrastructure on AWS.
  • ·         Deploy, manage, and troubleshoot containerized applications using Kubernetes.
  • ·         Develop automation tools, scripts, and operational utilities using Python.
  • ·         Build and maintain CI/CD pipelines for reliable and repeatable application deployments.
  • ·         Monitor system health, availability, performance, and capacity. Implement observability using metrics, logs, traces, dashboards, and alerting.
  • ·         Participate in incident response, troubleshooting, root-cause analysis, and post-incident reviews. Define and improve SLIs, SLOs, and SLAs. Automate repetitive operational tasks and reduce manual intervention.
  • ·         Perform Kubernetes troubleshooting, including pods, deployments, services, ingress, networking, storage, and resource management. Optimize AWS infrastructure for performance, reliability, security, and cost.
  • ·         Implement infrastructure as code using tools such as Terraform or CloudFormation.
  • ·         Establish and maintain backup, disaster recovery, and business continuity mechanisms.
  • ·         Work with development teams to improve application reliability and production readiness.
  • ·         Participate in an on-call rotation and respond to production incidents when required.
  • ·         Continuously identify opportunities to improve system resilience and operational processes.
  • ·         Skills: Digital : Python~Digital : Kubernetes~Digital : Site Reliability Engineering (SRE)~AWS DevOps and Automation
  • ·         Experience Required: 8-10
  • ·         ** All submissions must have LinkedIn id of Candidate**

Must Have Skills:

  • ·         Apple
  • ·         AWS
  • ·         Kubernetes
  • ·         Python
  • ·         SRE

 

Thanks and Regards,

Grace

Technical Recruiter |ASCII Group LLC. 

Email:  |Direct:

 

 

Similar jobs