Sr AIOps and Incident Management SRE
Quick Overview
Job Description
Job title: Senior AIOps and Incident Management / Site Reliability Engineering
Location: Hybrid Onsite (Fort Mill, SC, Austin TX, Boston, MA, New York, NY, Tempe, AZ, San Diego, CA)
Duration: Contract
Job Description:
We are seeking a Senior AIOps Incident Manager & Site Reliability Engineer to lead incident management, operational resilience, and intelligent automation initiatives across enterprise environments. This role will partner with Network Operations Center (NOC), Infrastructure Operations, Cloud Engineering, DevOps, and Application Support teams to proactively detect, respond to, and prevent technology incidents. The ideal candidate combines hands on incident management expertise with experience implementing observability, automation, and AI-driven solutions to improve system reliability, reduce operational overhead and enhance customer experience. The candidate should Possess expertise in AIOps, ITSM, ITIL, SRE, Incident Management, Cloud Operations, and Enterprise Infrastructure.
Incident & Recovery Management
- Monitor, document, and analyze major incident response efforts and service recovery activities.
- Serve as a senior escalation point for Tier 1 and Tier 2 operational incidents.
- Conduct incident reviews, root cause analysis, and corrective action planning.
- Improve Mean Tme to Detect (MTTD) and Mean Time to Resolve (MTTR).
Site Reliability Engineering
- Implement SRE practices to improve platform reliability, scalability, and resiliency.
- Define and monitor SLAS, SLOS, and operational KPIs.
- Develop proactive reliability and availability strategies.
AIOps & Automation
- Implement AIOps solutions to automate incident detection, remediation, and prevention.
- Build and optimize AI-powered Operational agents and self-healing workflows.
- Reduce operational effort through intelligent automation.
Observability & Monitoring
- Lead enterprise monitoring initiative using Dynatrace and related observability platforms.
- Improve visibility across cloud, infrastructure, applications, and user experiences.
- Enable predictive monitoring and anomaly detection.
ITSM & Service Operations
- Develop and enhance incident, problem, change, and event management frameworks aligned with ITIL and ITSM best practices.
- Leverage ServiceNow workflow automation to improve service delivery.
Cross-Functional Leadership
- Partner with Infrastructure, DevOps, Cloud, Security, Application Development, and NOC teams.
- Mentor operational teams and promote an automation-first culture.
Qualifications
- Bachelor''s degree in Computer Science, Information Engineering, or related field (or equivalent experience).
- 8+ years Of experience in IT Operations, Site Reliability Engineering, Infrastructure Operations, Network Operations, or Production Support environments.
- 5+ years of experience leading incident management, transformation, or reliability engineering
- Strong experience with:
- Site Reliability Engineering (SRE)
- IT Service Management (ITSM)
- ITIL Framework
- Incident, Problem, Change, and Event Management
- Network Operations Center (NOC)
- Infrastructure Operations
- Service Desk
- Application Support
- Cloud Platforms (AWS Azure, or Google Cloud Platform)
- DevOps Practices and Toolchains
- Hands on experience with Dynatrace, monitoring platforms, and observability solutions.
- Experience using ServiceNow for ticketing, workflow automation, and service management.
- Strong understanding Of infrastructure, networking, cloud architecture, and enterprise application ecosystems
- Proven experience conducting root cause analysis and implementing preventive controls.
- Experience leading enterprise AIOps implementations.
- Experience building AI-Powered operational agents and intelligent automation solutions.
- Certifications such as:
o ITIL Foundation or ITIL Managing Professional
o Certified Site Reliability Enginær (SRE)
o AWS Azure, or Cloud certifications
o ServiceNow certifications
- Experience with workflow orchestration and enterprise automation platforms.
- Familiarity with predictive analytics, machine learning operations, and autonomous operations frameworks.
Educational Qualifications:
- Required - Bachelor’s degree in Computer Science, Information Technology, Computer Engineering or closely related or equivalent.
- Preferred - Master’s degree in Management Information Systems (MIS), Computer Science, Big Data or Analytics or equivalent.
Travel:
· Open to travel based up on the nature of the engagement.
Thanks & Regards
Srikanth Donkani Resource Manager | Reliable Software Direct: |
AI & Analytics Generative AI Machine Learning Cloud DevOps SAP Data Engineering Data Science Databricks Snowflake |
Industries: Government | Healthcare | Banking | Manufacturing | Retail ISO Cert: 9001 | 27001
Equal Employment Opportunity Reliable Software employment does not discriminate on the basis of race, religion, gender, sexual orientation, age or any other basis as covered by federal, state, or local law. Employment decisions are based solely on qualifications, merit and business needs. |
Skills
Similar jobs
AWS Databricks DevOps Lead / Cloud Data Platform DevOps Lead
Montek System · Plano, United States
13 minutes agoHiring For "Senior Devops/Kubernetes Engineer" Bellevue, WA (hybrid)
Empower Professionals · Bellevue, United States
13 minutes ago€50/hrDevOps Engineer || Houston, TX (Onsite) Must be Local
AKAASA Technologies · Houston, United States
31 minutes agoPrincipal Platform Engineer/DevSecOps Architect - W2 Contract
Arrowminds inc · Irving, United States
32 minutes ago.Net Developer with DevSecOps principles, Azure DevOps with Entra ID
ARK Infotech Spectrum · Mahwah, United States
32 minutes agoSenior AWS Platform Engineer
AGM Tech Solutions, LLC · Alpharetta, United States
33 minutes ago