Haystack
← Back to Jobs
Technology

Site Reliability & Operations Engineer (L3 Support) - Blue Ash, OH (5 days onsite)

Activesoft, Inc.Blue Ash, OH🇺🇸United StatesPosted 23 Jul 2026

Why This Role Stands Out

This role offers a fantastic opportunity to lead critical incident response and drive system reliability within a reputable technology company, allowing you to hone your L3 support and RCA skills. If you thrive on solving complex technical challenges and collaborating with diverse teams to ensure seamless operations, you'll find significant growth and impact here. Apply to become an integral part of their dedicated operations team!

Quick Overview

Work Type
On Site
Level
Mid Senior

Job Description

Job Title: Site Reliability & Operations Engineer (L3 Support)
Location: Blue Ash, OH (5 days onsite)
Job Duration: 18 Month(s) - CTH
Travel: Required (approximately every 4 weeks; additional travel during new site go-lives)

Key Responsibilities:

  • Lead and manage Major Incident (P1/P2) response, including bridge/war-room coordination and stakeholder communication.
  • Drive Root Cause Analysis (RCA), Problem Management, and corrective actions through to closure.
  • Provide L3 production support for business-critical applications and infrastructure.
  • Monitor, troubleshoot, and resolve production issues across cloud, on-premises, and distributed environments.
  • Partner with development, infrastructure, and business teams to resolve incidents and improve platform reliability.
  • Participate in application design, testing, deployment, and software delivery improvements.
  • Support infrastructure and applications through a rotating 24x7 on-call schedule.
  • Create and maintain operational procedures, documentation, and support processes.

Required Skills:

  • Hands-on experience leading Major Incident Management (P1/P2) from start to resolution.
  • Strong experience with Root Cause Analysis (RCA) and Problem Management methodologies (5 Whys, Fishbone, Timeline Analysis, etc.).
  • Experience providing L3 Production Support or Site Reliability Engineering (SRE) support in enterprise environments.
  • Strong troubleshooting skills using monitoring tools such as Dynatrace, Splunk, Grafana, AppDynamics, or similar.
  • Experience supporting distributed applications across cloud and on-premises environments.
  • Excellent communication and stakeholder management skills.

Skills

Splunk
Grafana
Stakeholder Management

Similar jobs