Haystack
← Back to Jobs
Engineering
GT

Lead Site Reliabitility Engineer

GSPANN TechnologiesSt. Louis, MO🇺🇸United StatesPosted Oct 1, 2026

Why This Role Stands Out

This Lead Site Reliability Engineer role at GSPANN Technologies offers a fantastic opportunity to enhance the reliability and performance of global e-commerce platforms, driving innovation in observability solutions. You'll thrive here if you're passionate about building resilient systems and eager to lead best practices in a dynamic, client-focused environment. Apply today to make a significant impact within a reputable IT services firm.

Quick Overview

Seniority
Mid Senior
Work mode
On Site
Location
St. Louis, MO, United States
Posted
5 days ago

Job Description

About GSPANN

Headquartered in California, U.S.A., GSPANN provides consulting and IT services to global clients. We help clients transform how they deliver business value by helping them optimize their IT capabilities, practices, and. With five global delivery centers and 2000+ employees, we provide the intimacy of a boutique consultancy with the capabilities of a large IT services firm.

Role: Lead Site Reliability Engineer
Location: St Louis, MO (5 days Onsite)
Duration: 12+ months

The Site Reliability Engineer will run the production environment by monitoring availability and taking a holistic view of system health. They will build software and systems to manage platform infrastructure and applications; improve reliability, quality, and time-to-market of our suite of software solutions; and measure and optimize system performance - all with an eye toward pushing our capabilities forward, getting ahead of customer needs, and innovating to continually improve.

Responsibilities

  • Ensure availability, latency, performance, and efficiency of our global ecommerce sites
  • Experience driving change management and incident management
  • Promote best practices and innovative observability to guide product delivery teams in achieving operational excellence for new product deliveries.
  • Drive operational excellence and evangelize best practices in observability.
  • Develop unified observability dashboards and implement E2E observability requirements.
  • Design innovative observability solutions for internal and external stakeholders.
  • Contribute to observability instrumentation standards and create repeatable patterns for engineering teams.
  • Define and implement E2E observability requirements and lead teams to support E2E best practices.
  • Collaborate with cross-functional teams to achieve objectives and drive high reliability into systems.
  • Build proprietary tools to mitigate weaknesses in incident management or software delivery.
  • Implement SRE best practices to increase system reliability and performance.
  • Automate processes for improved collaborative response and prepare teams for incidents.
  • Maintain error budgets, meet SLOs, and support uptime and availability of critical platform components.
  • Automate technology stacks to improve operating costs while responding to traffic spikes.

Required Skills and Experience:

  • Bachelor's Degree in Computer Science, Information Science, Engineering, or a related field.
  • 10+ years of experience in code management, deployment processes, procedures, and tools in a DevOps or SRE role.
  • Experience with monitoring tools (preferred: Dynatrace, Splunk, Datadog, Grafana, and New Relic).
  • Proficiency in state-of-the-art observability trends, tools, products, and technologies.
  • Ability to identify organization-wide gaps in the SRE practice and implement solutions that contribute to organizational transformation.
  • Experience driving cross-organization adoption of new technologies or initiatives.
  • Ability to influence senior management in selecting the right strategy, processes, and structures to transform the organization into a modern SRE team.
  • Proactive in identifying performance bottlenecks, anomalous system behavior, and addressing root causes of service issues.
  • Passionate about technology with a strong sense of curiosity and a desire to improve processes, automate everything, and continuously learn.
  • Successful experience supporting a cloud production environment (strong preference for Azure).
  • Competency in one or more programming languages for automation (Python strongly preferred).
  • Knowledge of cloud deployment tools and methodologies (ideally Ansible, but Terraform, Azure DevOps, etc. are also considered).
  • Deep understanding of Kubernetes and Docker architecture and associated tools.
  • Experience with at least one configuration management solution (e.g., Chef, Ansible, AWS CodeDeploy).
  • Proficiency with repository and pipeline-related tools (e.g., GitLab, Jenkins, Bamboo, Travis, CircleCI).
  • Experience with implementing and using various application and infrastructure monitoring tools.
  • Strong troubleshooting skills.
  • Ability to take ownership and deliver solutions autonomously.

Working at GSPANN

GSPANN is a diverse, prosperous, and rewarding place to work. We provide competitive benefits, educational assistance, and career growth opportunities to our employees. Every employee is valued for their talent and contribution. Working with us will give you an opportunity to work globally with some of the best brands in the industry.

The company does and will take affirmative action to employ and advance in the employment of individuals with disabilities and protected veterans and to treat qualified individuals without discrimination based on their physical or mental disability status. GSPANN is an equal opportunity employer for minorities/females/veterans/disabled.

Similar jobs