Quick Overview
Job Description
JD:
Education: Master s degree preferred Bachelor s degree in Computer Science, Information Technology, Systems
Engineering, or a related field required.
Experience: 8+ years leading complex technical projects, product delivery, or cloud infrastructure rollouts within a global
enterprise IT environment, with at least 2+ years leading AIAIOps platform deployments.
Deep AI & Agentic Literacy: Mastery of AI concepts, including Large Language Models (LLMs), Agentic AI, multi-agent
reasoning, tool-calling architectures, Retrieval-Augmented Generation (RAG), and the Model Context Protocol (MCP).
Observability & Cloud Baseline: Hands-on experience with enterprise monitoring platforms (e.g., Dynatrace, Splunk,
CloudWatch), cloud environments (Google Cloud Platform, AWS), and CICD pipeline automation.
Systems Thinking & Governance: Proven ability to translate complex I&O environments into contextual schemas for AI
consumption, paired with experience establishing guardrails for data security, AI safety, and compliance.
SRE & DevOps Mindset: Deep familiarity with Site Reliability Engineering (SRE) principles, automated testing, IaC, and
data pipeline oversight (e.g., SnowflakeETL validation).
Roles & Responsibilities
In this role, you will serve as the primary operational driver and technical delivery lead across key automation initiatives including
agentic AI platforms, enterprise observability frameworks, vulnerability tracking, and automated self-healing solutions. Operating
with a high degree of autonomy, you will bridge architectural strategy with hands-on technical delivery, coordinating cross-functional
engineering teams to ensure milestones are met, financial savings are realized, and executive leadership receives strategic
visibility.
Key Responsibilities
Orchestrate Agentic AI & AIOps Delivery: Serve as the technical delivery orchestrator for enterprise-wide Agentic
AIOps initiatives. Drive milestone tracking, integration of developed code blocks, and deployment of multi-agent
reasoning systems, dynamic agent registries, and tool-calling interfaces utilizing Google Cloud Vertex AI, Model Context
Protocol (MCP), and agentic frameworks.
Lead Observability & Self-Healing Automation: Direct operational delivery for enterprise observability programs.
Oversee telemetry ingestion, monitoring standardization, automated alerting, and self-healing automation workflows
across multi-cloud infrastructure using enterprise observability tools (e.g., Dynatrace, Splunk, CloudWatch).
Vulnerability Tracking & Security Remediation: Formulate, integrate, and deploy standardized vulnerability tracking workflows and automated remediation frameworks across application and infrastructure portfolios to maintain strict enterprise compliance and security posture. Cross-Functional Team Orchestration: Harmonize diverse engineering, application support, data, and security teams. Manage resource dependencies, navigate competing priorities, and drive strict accountability across separate technical units to meet consolidated timeline goals. Monthly Project Reporting & Savings Tracking: Establish quantitative KPIs to track delivery health, token consumption, operational MTTR improvements, and financial value. Build and maintain executive reporting frameworks that explicitly measure and report cost savings generated by AI and automation initiatives. Executive Communication & Autonomy: Establish concise, high-level reporting cadences for executive leadership. Anticipate bottlenecks, clear technical roadblocks independently, and distill complex, multi-team technical statuses into clear, actionable executive updates.
Generic Managerial Skills, If any
Should be able to work with Client directly as per the requirements at Onsite
Client engagement, leadership, global delivery, stakeholder management, thought leadership,
reporting & governance
Strong leadership and decision-making abilities
Excellent communication and executive presence
Ability to manage complex, high-visibility programs
Experience working in global delivery models
Key Words to search in Resume
Google Cloud Platform, AIOps, Automation, GCVE, Hybrid Cloud, AIML, LLM, Power Shell, Ansible, Shell Script
Pre-Screening Questionnaire
Explain an Agentic AI architecture you have designed or implemented. What problem did it solve?
What is the difference between traditional AI workflows and Agentic AI systems?
How have you used LLMs in production environments?
What is RAG (Retrieval-Augmented Generation), and where have you applied it?
What experience do you have with MCP (Model Context Protocol) or tool-calling architectures?
Role Descriptions: AI Ops lead
Essential Skills: AI Ops lead
Desirable Skills:
Keyword:
Skills: Digital : AnsibleDigital: Terraform
Experience Required: 8-10
Similar jobs
- LE
Senior Product Manager - Enterprise
NewLegora
New York City🇺🇸4 hours agoOperations & Project Management - IN
Regional Director (Continuous Opening)
NewInstructure
US-REMOTE🇺🇸Remote11 hours agoSalesforceCRMData Privacy+3Operations & Project Management - IE
Senior Learning Program Manager
NewIndustrial Electric Manufacturing
San Antonio🇺🇸2 hours agoArticulateComplianceContinuous Improvement+9Operations & Project Management - FL
Project Manager, Customer Success
NewFlaglerhealth
NYC Office🇺🇸Remote2 hours agoEHRTriageOperations & Project Management - CY
Associate Business Analyst - US
NewCytora.com
Eastern🇺🇸Remote2 hours agoAgileBusiness AnalysisCompliance+5Operations & Project Management - EQ
General Manager (Site Solutions)
EquipmentShare
Atlanta🇺🇸$80k - $100k/yr3 days agoContinuous ImprovementOperations & Project Management