Quick Overview
Seniority
Mid Senior
Work mode
Hybrid
Location
Houston, TX, United States
Posted
19 hours ago
MicroservicesAnsibleAzureBash.NETJavaKubernetesLESSPowerShellPythonTerraform
Job Description
Title: Site Reliability Engineer
Location: Houston, TX 77002 (Hybrid: 3 days onsite / 2 days remote)
Duration: Contract to Hire
Work Requirements: U.S. Citizen, Holders, or Authorized to Work in the U.S.
Job Description:
The Site Reliability Engineer is a founding member of SRE practice's.
This role exists to move the organization from reactive operations to engineered reliability. You will study how our most critical systems fail, particularly our internal applications and facility automation interfaces and design controls, automation, and observability that reduce incidents over time.
Success in this role means fewer false alerts, faster recovery, less manual intervention, and systems that heal themselves when possible.
You will work closely with application, infrastructure, and operations teams and participate directly in on call and incident response.
What You Will Own
Technical Environment
About INSPYR Solutions
Technology is our focus and quality is our commitment. As a national expert in delivering flexible technology and talent solutions, we strategically align industry and technical expertise with our clients' business objectives and cultural needs. Our solutions are tailored to each client and include a wide variety of professional services, project, and talent solutions. By always striving for excellence and focusing on the human aspect of our business, we work seamlessly with our talent and clients to match the right solutions to the right opportunities. Learn more about us at inspyrsolutions.com.
Location: Houston, TX 77002 (Hybrid: 3 days onsite / 2 days remote)
Duration: Contract to Hire
Work Requirements: U.S. Citizen, Holders, or Authorized to Work in the U.S.
Job Description:
The Site Reliability Engineer is a founding member of SRE practice's.
This role exists to move the organization from reactive operations to engineered reliability. You will study how our most critical systems fail, particularly our internal applications and facility automation interfaces and design controls, automation, and observability that reduce incidents over time.
Success in this role means fewer false alerts, faster recovery, less manual intervention, and systems that heal themselves when possible.
You will work closely with application, infrastructure, and operations teams and participate directly in on call and incident response.
What You Will Own
- Definition and implementation of SLIs and SLOs that measure meaningful system health, not just availability
- Observability across the full stack, correlating cloud services, APIs, and on premise facility operations
- Automation to eliminate operational toil, including patching, data corrections, restarts, and recovery tasks
- Development of self healing behaviors for common failure modes
- Participation in on call rotations and leadership of blameless post incident reviews
- Design and execution of disaster recovery tests across SaaS, cloud, and on premise environments
Technical Environment
- Hybrid environments spanning cloud and on-premise infrastructure
- Azure cloud services
- Software Development/OOP skills within either .NET/Java/Python
- Observability tooling across logs, metrics, and alerting
- Automation using Python, PowerShell, Bash, or Ansible
- CI/CD tools and modern deployment practices
- Exposure to containerized and distributed systems environments
- 4+ years of experience in SRE, DevOps, Systems Engineering, or related roles
- Strong Linux and Windows systems administration and troubleshooting skills
- Hands-on experience with automation and scripting
- Experience designing and operating monitoring, alerting, and observability solutions
- Practical experience working in Azure environments
- Strong analytical skills and a bias toward eliminating root causes, not symptoms
- Ability to collaborate across application, infrastructure, and operations teams
- Exposure to Kubernetes, microservices, or container orchestration
- Hands-on experience with infrastructure as code tools such as Terraform or Ansible
- Understanding of distributed systems and high availability design
- Experience with SRE practices such as SLO based operations, runbook automation, or chaos testing
Our benefits package includes:
- Comprehensive medical benefits
- Competitive pay
- 401(k) retirement plan
- …and much more!
About INSPYR Solutions
Technology is our focus and quality is our commitment. As a national expert in delivering flexible technology and talent solutions, we strategically align industry and technical expertise with our clients' business objectives and cultural needs. Our solutions are tailored to each client and include a wide variety of professional services, project, and talent solutions. By always striving for excellence and focusing on the human aspect of our business, we work seamlessly with our talent and clients to match the right solutions to the right opportunities. Learn more about us at inspyrsolutions.com.
Similar jobs
- BO
Site Reliability Engineer (Associate, Experienced, or Senior)
NewBoeing
Saint Louis, Missouri🇺🇸$99.5k - $134.6k/yrOn-site9 minutes agoSQLAWSSonarQube+16Technology - ZC
Collibra DQ Platform Engineer
NewZtek Consulting
Charlotte, NC🇺🇸Hybrid19 hours agoSQLOAuthGit+4Technology - KT
Senior DevOps Engineer
NewKforce Technology Staffing
New York, NY🇺🇸Hybrid19 hours agoAzureBashKubernetes+4Technology - SS
W2 Role - SRE Production Support Engineer - Austin, TX (Onsite)
NewSDH Systems
Austin, TX🇺🇸On-site19 hours agoManufacturing - IU
AI Platform Engineer (Shared AI Capabilities)
NewIBOTIX US Inc.
Fort Worth, TX🇺🇸On-site19 hours agoAWSAssemblyAzure+3Technology - WH
DevSecOps Engineer
NewWhiztek Corp
Schaumburg, IL🇺🇸Hybrid19 hours agoGoogle WorkspaceEngineering