Why This Role Stands Out
Lead critical enterprise AI applications and drive continuous service improvement in a dynamic production environment, making a significant impact on platform stability and performance. This role is ideal for an experienced Production Application Support professional who thrives on troubleshooting, RCA, and proactive monitoring, offering a fantastic opportunity to shape the future of DevOps at a reputable organization. Apply today to leverage your expertise and advance your career in this impactful leadership position.
Quick Overview
Job Description
hackajob is partnering directly with BNY Mellon to hire for this role.
Were seeking a future team member for the role ofSenior Vice President, DevOps Production Services to join our team. This role is located in Manchester.
Role Overview:
We are seeking a highly skilled professional with strong experience in Production Application Support to manage and support critical enterprise AI based applications in a fast-paced production environment. The role requires hands-on expertise in monitoring, incident management, troubleshooting, release support, and ensuring high availability and stability of business-critical platforms.
In this role, youll make an impact in the following ways:
Provide L2/L3 production support for enterprise applications and ensure platform stability, resiliency, and availability.
Monitor application health, system performance, batch jobs, interfaces, and alerts using enterprise monitoring and observability tools.
Investigate, troubleshoot, and resolve production incidents within defined SLAs.
Perform root cause analysis (RCA) for recurring issues and drive permanent fixes.
Analyze production logs, identify failure patterns, and create actionable dashboards to improve service monitoring and incident response.
Coordinate with development, infrastructure, database, network, and business teams for issue resolution.
Support application deployments, change requests, weekend releases, and post-release validations.
Maintain incident, problem, and change records in service management tools.
Drive continuous service improvement through automation, process optimization, and proactive monitoring.
Participate in on-call support and major incident management as required.
Prepare operational reports, service health summaries, and stakeholder communications.
Write and analyze SQL queries for data validation, issue investigation, and production troubleshooting.
Use Unix/Linux commands and scripting for application support, log reviews, file handling, and system-level troubleshooting.
Leverage Splunk extensively for log analysis, issue diagnosis, trend identification, alerting insights, and dashboard creation.
To be successful in this role, were seeking the following:
Proven experience in production application support for business-critical applications.
Strong understanding of incident management, problem management, and change management processes.
Strong SQL skills for querying, troubleshooting, and data analysis in production environments.
Extensive hands-on experience with Splunk for log analysis, search creation, troubleshooting, monitoring, and dashboard development.
Strong Unix/Linux skills for navigating servers, reviewing logs, troubleshooting jobs/processes, and supporting application runtime environments.
Experience with monitoring and alerting tools, log analysis, Grafana, and dashboard-based production support.
Experience with ITSM tools such as ServiceNow, Jira, or similar platforms.
Ability to analyze application, infrastructure, and integration issues across distributed systems.
Experience supporting applications in cloud and/or on-prem environments.
Familiarity with scripting and troubleshooting middleware/interfaces.
Strong knowledge of release support, service recovery, and operational governance.
Ability to work in a high-pressure environment with strong ownership and accountability.
Preferred Skills
Azure Cloud experience preferred.
Knowledge of automation/scripting using Python, Shell, or PowerShell.
Exposure to DevOps / SRE practices, CI/CD pipelines, and observability tooling.
Strong communication skills with the ability to provide concise incident and executive status updates.
At BNY, our culture allows us to run our company better and enables employees growth and success. As a leading global financial services company at the heart of the global financial system, we influence nearly 20% of the worlds investible assets. Every day, our teams harness cutting-edge AI and breakthrough technologies to collaborate with clients, driving transformative solutions that redefine industries and uplift communities worldwide.
Recognized as a top destination for innovators, BNY is where bold ideas meet advanced technology and exceptional talent. Together, we power the future of finance and this is what #LifeAtBNY is all about. Join us and be part of something extraordinary.
Similar jobs
- TK
Platform Engineer
NewThe Key Support Services
United Kingdom🇬🇧Remote46 minutes agoGCPMongoDBCompliance+4Technology - OI
Senior Platform Engineer
NewOcean Infinity
London🇬🇧Hybrid2 hours agoRoboticsAzureKubernetes+2Technology - HA
Site Reliability Engineer / Production Support
NewHackajob Ltd
South East London, London🇬🇧Hybrid14 hours agoTechnology - VI
Platform Engineer
NewVIQU IT Recruitment
Bishops Itchington, Southam🇬🇧Hybrid14 hours agoAWSActive DirectoryAzure+3Technology - AM
Site Reliability Engineer
NewAnson Mccade
Manchester, Greater Manchester🇬🇧£40k/yrHybrid14 hours agoDockerMicroservicesMongoDB+15Technology - HA
Lead SRE - Chase UK
Hackajob Ltd
Charing Cross, Central London🇬🇧Hybrid3 weeks agoMicroservicesAWSLoad Balancing+6Technology