Haystack
← Back to Jobs
Technology
SM

Investment Banking Cloud Engineer - NYC Onsite - 12+ yrs

Systems Management Group, IncNew York, NY🇺🇸United StatesPosted Aug 23, 2026

Why This Role Stands Out

Leverage your extensive cloud engineering expertise to enhance critical financial risk systems, driving operational excellence and automation within a reputable firm. This role offers a fantastic opportunity to contribute directly to the stability and advancement of core trading technologies, perfect for a proactive problem-solver eager to make a significant impact. Apply today to join a collaborative team and advance your career in a dynamic on-site environment.

Quick Overview

Seniority
Mid Senior
Work mode
On Site
Location
New York, NY, United States
Posted
3 weeks ago
SQLETLSnowflakeAgileAngularAzureConfluenceDatabricksJavaJiraLESSPython

Job Description

Role: Investment Banking Cloud Engineer

Exp: 12+ yrs

Location: New York, NY (Onsite)

Interview: Final Round In-Person in NYC

Job Summary:

  • We are is seeking a Investment Banking Cloud Engineer consultant to support production stability, operational execution, and continuous improvement across CCR technology platforms.
  • The successful candidate will work with the CCR Operations IT Lead, onsite team members, offshore resources, engineering teams, and business stakeholders to help operate and improve a cloud-enabled IT operations model while reducing manual effort and operational risk through automation and process discipline.
  • The candidate should be a hands-on cloud engineer with strong production support discipline, the ability to troubleshoot issues, communicate clearly, manage multiple tasks, and contribute effectively within a collaborative team environment.

Major Responsibilities:

Exposure Management Application Support:

  • Support key production operations for applications supporting EM, including EPE/PFE calculations, limit monitoring, what-if pre-trade intraday analysis, and reporting.
  • Ensure platforms consistently meet the availability, performance/SLAs, and data-quality expectations of EM users during daily, intraday, and regulatory cycles (monthly, quarterly, annually).
  • Coordinate with onsite and offshore IT members to translate business priorities into day-to-day operational execution.

Production Operations & Continuous Improvement:

  • Demonstrate a continuous improvement mindset through their example, with a focus on reducing incidents, manual interventions, and operational risk, and improving turnaround time for incidents.
  • Contribute to measurable operational improvements using data points and KPIs/metrics such as availability, incident trends, MTTR, SLA breaches, reruns, and batch success rates.
  • Identify and execute practical opportunities for process standardization, tooling enhancements, and Operational simplification across the CCR application stack.

Global Delivery & Offshore Leverage:

  • Support a globally distributed operating model, leveraging offshore teams for L1/L2 support, batch operations oversight, upstream data feed validation, data quality monitoring.
  • Follow and help refine clear onshore/offshore operating boundaries, escalation paths, and ownership models to support seamless production coverage.
  • Contribute knowledge transfer, documentation standards, and runbook maturity to improve offshore effectiveness and reduce dependency on key individuals, ensuring service quality, productivity, and continuous skill uplift.

Incident, Problem & Change Management:

  • Timely escalate highseverity production incidents, particularly those impacting exposure reporting, limit breaches, what-if intraday analysis or regulatory deliverables.
  • Perform root cause analysis (RCA) for impactful incidents (Calculation breaks, data validation/reconciliation issues), ensuring recurring issues are eliminated through permanent fixes rather than shortterm workarounds.
  • Support change and release governance aligned with Client CAB, DevOps, CI/CD, risk, audit, and control standards.

Platform Reliability, Automation & Resilience:

  • Implement automation and self healing for batch monitoring, data validation, reconciliations, and recovery processes for production as well as lower testing environments.
  • Drive with Engineering and Infrastructure teams to improve observability, alerting, capacity planning, and resilience of CCR platforms.
  • Ensure DR/BCP readiness for EM critical systems, including regular/annual testing and documented recovery procedures.

Regulatory, Audit & Risk Controls:

  • Ensure CCR applications meet regulatory, audit, security, and internal risk management requirements.
  • Support IT Operations activities related to internal audits, regulatory exams, and model governance reviews by providing evidence, documentation, and control execution support.
  • Maintain robust IT control frameworks (Access management, change control, data integrity).

Required Qualifications:

Experience:

  • 10+ years of experience in IT Operations / Production Support within financial services.
  • Experience a senior consultant role supporting risk, exposure, or comparable financial technology platforms with hands-on work capabilities.
  • Demonstrated experience operating the offshore leveraged global support models in a regulated environment.
  • Bachelor s Degree in Computer Science, Management of Information Systems, or related business discipline(s).

Technical & Domain Expertise:

  • Understanding of Exposure Management and Investment Banking concepts, including derivative trade cycles, market data, PFE, EPE, Collateral/margin management, limits, stress testing, and whatif analysis.
  • Experience supporting batch intensive and intraday realtime risk platforms (Java, Angular, Python, PySpark, ETL, Stored Procedure, SQL, Snowflake, PowerBI).
  • Hands-on experience with Cloud computing and infrastructure (Databricks, Medallion architecture, Azure, ADF, Spark based distributed compute, other Cloud native technologies).
  • Proven track record driving operational excellence, automation, and process maturity with an Agile based application/software development.
  • Proven ability to utilize ServiceNow, JIRA, Confluence, PowerBI to manage incidents, tasks, and releases, generate a system diagnosis report, and demonstrate KPIs based operational improvement.
  • Familiarity with ITIL based service management (ITSM) frameworks, DevOps, CI/CD, and reliability engineering practices.

Leadership & Soft Skills:

  • Continuous improvement mindset, with the ability to contribute to operational maturity across onshore and offshore teams.
  • Strong experience collaborating with distributed, multicultural teams and thirdparty vendors.
  • Excellent interpersonal, written and verbal communication skills with the ability to engage Exposure Management users, and senior Technology team leads.
  • Ability to prioritize and effectively manage multiple tasks under time-critical and regulatory pressure.
  • Highly motivated, self-directed individual with the ability to work independently and in team environments.
  • Effective presenter capable of articulate strategic execution plans including system/architecture/data flow diagrams, combined with extreme attention to detail.
  • Strong analytical, problem solving (breaking ambiguolarge problems to smaller/less complex ones), and decision-making skills.
  • Collaborative team player and relationship builder.

 

Regards,

Prakash

prakash.v(@)smg(-)llc(.)us

Similar jobs