Haystack
← Back to Jobs
Engineering

Senior IT Triage Engineer

First-Citizens Bank & Trust CompanyPhoenix, AZ🇺🇸United StatesPosted 23 Jul 2026

Why This Role Stands Out

This hybrid role at First-Citizens Bank offers a fantastic opportunity to lead critical incident resolution and drive continuous improvement within a reputable financial institution, allowing you to significantly impact enterprise-wide service stability. You'll thrive here if you possess strong troubleshooting skills and a passion for problem-solving in high-pressure environments, making this an excellent step for your career growth. Apply today to join a team dedicated to operational excellence and technological advancement.

Quick Overview

Work Type
Hybrid
Level
Mid Senior

Job Description

Overview

The Senior IT Triage Engineer serves as a critical technical leader responsible for the rapid diagnosis, restoration, and resolution of complex technology incidents across enterprise infrastructure, cloud platforms, applications, and end-user services. This role combines deep technical troubleshooting expertise with a proactive focus on eliminating recurring issues through structured problem management, root cause analysis, and continuous service improvement.

The ideal candidate excels in high-pressure operational environments, drives technical resolution efforts across multiple teams, and leverages automation, observability, and reliability practices to improve service stability and reduce operational risk.

Responsibilities

Key Responsibilities
  • Incident Management, Service Restoration, and System Analysis.
  • Lead technical triage efforts for high-priority incidents and service disruptions.
  • Coordinate cross-functional teams to restore critical business services as quickly as possible.
  • Analyze alerts, logs, monitoring data, and telemetry to identify the source of issues.
  • Serve as a senior escalation point for complex infrastructure, cloud, network, application, and platform incidents.
  • Drive incident bridges, facilitate technical discussions, and maintain clear communication with stakeholders.
  • Problem Management & Root Cause Elimination
  • Lead root cause investigations for recurring or significant incidents.
  • Develop and drive corrective and preventive action plans across technology teams.
  • Track and manage problem records through resolution.
  • Identify systemic issues and technical debt that impact service stability.
  • Attend post-incident reviews and ensure lessons learned are translated into operational improvements.

Reliability & Operational Excellence

  • Continuously improve service availability, resiliency, and operational performance.
  • Partner with engineering teams to improve monitoring, alerting, telemetry, and observability capabilities.
  • Reduce alert noise through event correlation, automation, and process optimization.
  • Drive efforts to improve mean time to detect (MTTD) and mean time to restore service (MTTR).

Automation & Process Improvement
  • Identify opportunities to automate operational workflows and repetitive support activities.
  • Develop runbooks, playbooks, and operational procedures.
  • Collaborate with engineering teams to implement self-healing and automated recovery mechanisms.

Technical Leadership

  • Provide mentorship and guidance to engineers and operational support teams.
  • Influence technical decision-making related to operational readiness and supportability.
  • Act as a trusted advisor for service reliability, supportability, and operational risk management.
  • Participate in change reviews to ensure production readiness and minimize operational impact.

Qualifications

Bachelor's Degree and 8 years of experience in Technical work in Application Development, Server Administration, Information Security, or Engineering OR High School Diploma or GED and 12 years of experience in Technical work in Application Development, Server Administration, Information Security, or Engineering
  • 7+ years of experience in enterprise IT operations, infrastructure engineering, platform operations, or technical support environments.
  • Proven experience managing and resolving critical production incidents across multiple system architectures and infrastructures.
  • Strong background in problem management and root cause analysis methodologies.
  • Experience supporting large-scale enterprise environments.
  • Strong understanding of ITIL service management practices.

Technical Skills
  • Modern application architecture patterns and operations
  • Infrastructure & Platforms
  • Windows and Linux administration
  • Virtualization platforms (VMware, Hyper-V, etc.)
  • Storage and backup technologies
  • Containers and orchestration platforms (OpenShift)

Monitoring & Observability

  • Enterprise monitoring platforms
  • Log aggregation and analysis tools
  • Application performance monitoring (APM)
  • Operational dashboards and telemetry systems
  • Event management solutions

Automation
  • PowerShell, Python, Bash, or equivalent scripting
  • Workflow automation platforms
  • API integration and orchestration
  • Runbook development

Key Competencies
  • Exceptional troubleshooting and analytical skills.
  • Strong sense of ownership and accountability.
  • Ability to perform effectively during major incidents and high-pressure situations.
  • Excellent collaboration and stakeholder management skills.
  • Strong written and verbal communication.
  • Data-driven decision making.
  • Systems thinking and problem-solving mindset.
  • Continuous improvement orientation.

Benefits are an integral part of total rewards and First Citizens Bank is committed to providing a competitive, thoughtfully designed and quality benefits program to meet the needs of our associates. More information can be found at

$descr2

$descr3

Skills

Stakeholder Management

Similar jobs