Haystack
← Back to Jobs
Technology
CM

Sr. Manager, Site Reliability Engineer

Credence Management SolutionsHolmdel, NJ🇺🇸United StatesPosted Sep 19, 2026

Why This Role Stands Out

This Sr. Manager, Site Reliability Engineer role offers a competitive salary and the opportunity to lead a global SRE team, advancing a modern operating model and significantly impacting system reliability. You'll thrive here if you possess strong technical leadership skills in SRE, cloud architecture, and automation, and you'll benefit from a hybrid work environment that fosters collaboration and professional growth. Apply now to shape the future of reliability at a leading enterprise hiring platform.

Quick Overview

Salary
$150k - $170k/yr
Seniority
Mid Senior
Work mode
Hybrid
Location
Holmdel, NJ, United States
Posted
16 hours ago
DynamoDBAWSNew RelicGrafanaTerraform

Job Description

Job Overview
We are seeking a Sr Manager of Site Reliability Engineering (SRE) to lead and continue developing our global SRE organization across the US, Ireland, India, and strategic engineering partners.
This role is responsible for advancing a modern SRE operating model that combines centralized reliability capabilities with product-aligned SRE support. The Sr Manager will partner closely with Product Engineering, Cloud Engineering, Operations, Security, DBA, and other technical teams to improve reliability, strengthen operational practices, and help development teams implement consistent engineering standards.
This is a highly technical leadership role requiring strong experience across Site Reliability Engineering, AWS cloud architecture, observability, cloud engineering, automation, incident and problem management, and FinOps. The successful candidate will be able to evaluate technical decisions across reliability, scalability, security, operational complexity, and cloud financial impact .
About Us
ICIMS is a leading enterprise hiring platform that combines the scale and reliability of enterprise software with the transformative power of AI. Thousands of organizations across more than 200 countries and territories trust ICIMS to find and hire the people who shape their future and drive their business forward. Powered by insights from billions of hiring interactions, continuous AI innovation, and a highly extensible platform, ICIMS helps organizations turn talent acquisition into a competitive advantage. For more than 25 years, ICIMS has delivered end-to-end hiring solutions that improve recruiting efficiency, reduce costs and create exceptional candidate experiences.
ICIMS helps solve one of the biggest challenges businesses face today: building a workforce that can adapt, scale, and perform in an increasingly competitive and unpredictable talent market. We uniquely do that by combining enterprise-grade hiring technology, AI-powered insights and automation, and connected talent experiences to help organizations improve hiring outcomes while driving measurable impact.
Responsibilities
Leadership & Strategy
  • Lead and develop a globally distributed SRE organization across multiple regions and time zones.
  • Define and execute the SRE strategy, operating model, priorities, and technical direction.
  • Establish a hybrid SRE model combining centralized Reliability Enablement capabilities with product-aligned SRE support.
  • Define clear responsibilities and decision rights across SRE, Product Engineering, Operations, Cloud Engineering, DBA, and other technical functions.
  • Build strong technical leadership, product ownership, regional handoffs, and knowledge-sharing practices across the global organization.
  • Develop engineers and technical leaders through coaching, mentorship, and clear technical and career expectations.
  • Drive a culture focused on proactive reliability engineering rather than reactive operational support.

Reliability Engineering & Product Alignment
  • Partner with Product Engineering teams to understand product architecture, service dependencies, reliability risks, and operational requirements.
  • Establish and mature SLIs, SLOs, error budgets, service-health measures, and operational-readiness standards where appropriate .
  • Identify systemic reliability issues before they become customer-impacting incidents.
  • Translate production learnings, recurring failures, and RCAs into prioritized engineering improvements.
  • Establish reusable reliability patterns, tooling, automation, runbooks, and engineering practices that can be adopted across product teams.
  • Ensure SRE supports product teams without replacing Product Engineering ownership of application functionality and defects.

Problem Management & Operational Excellence
  • Provide technical leadership during significant and complex production incidents.
  • Partner with Operations and engineering teams to strengthen incident response, escalation, restoration, and recovery practices.
  • Lead the continued development of structured problem management and root-cause analysis practices.
  • Connect recurring incidents and operational risks to visible, prioritized corrective actions.
  • Improve operational readiness, capacity planning, performance management, and service resiliency.
  • Reduce repetitive operational work and manual intervention through automation and engineering.

Observability & Reliability Enablement
  • Lead the development and adoption of common observability standards across logging, metrics, tracing, dashboards, monitoring, and alerting.
  • Drive enterprise observability strategy and governance across platforms including Grafana, OpenTelemetry , Sumo Logic, New Relic, CloudWatch, and related technologies.
  • Establish practical service blueprints and reusable observability patterns for development teams.
  • Improve alert quality, service visibility, dependency awareness, and actionable monitoring.
  • Ensure observability capabilities support both real-time incident response and longer-term reliability improvement.

Operational FinOps & Cloud Financial Management
  • Incorporate cloud financial awareness into architecture, reliability, and engineering decisions.
  • Partner with Cloud Engineering, Finance, Product, and engineering leadership to improve cloud cost ownership and accountability.
  • Identify opportunities for rightsizing, workload optimization, storage efficiency, and removal of unnecessary cloud consumption.
  • Understand AWS pricing and commitment constructs including Savings Plans, Reserved Instances, licensing considerations, and consumption-based services.
  • Support effective cost allocation, tagging, forecasting, reporting, and cloud financial governance.
  • Evaluate technical decisions using both engineering and financial considerations while ensuring cost optimization does not introduce unacceptable reliability or performance risk.

AWS & Cloud E ngineering
  • Provide senior technical leadership for complex cloud environments, with AWS as the primary platform.
  • Partner with Cloud Engineering on architecture, resiliency, automation, networking, security, governance, and platform standards.
  • Review and guide architectures involving AWS technologies such as ECS, ECR, EC2, RDS, Aurora, S3, DynamoDB, OpenSearch, SQS, SNS, Kinesis, IAM, AWS Organizations, and cloud networking.
  • Evaluate architecture across availability, scalability, performance, security, recoverability, operational complexity, and cost.
  • Drive Infrastructure as Code, automated provisioning, standardized cloud patterns, and policy-based governance

Qualifications
  • 10+ years of experience across Site Reliability Engineering, Cloud Engineering, Platform Engineering, DevOps, Infrastructure Engineering, or related technical disciplines.
  • 5+ years of technical or engineering leadership experience, including responsibility for technical strategy, team leadership, and organizational outcomes.
  • Experience leading and collaborating with distributed technical teams across multiple regions and time zones.
  • Strong technical knowledge of AWS and experience supporting or designing complex enterprise cloud environments.
  • Broad understanding of cloud computing , containers, networking, storage, databases, security, identity, monitoring, and governance.
  • Strong understanding of highly available , scalable, resilient, and distributed production systems.
  • Experience with Infrastructure as Code and automated infrastructure delivery, preferably Terraform.
  • Strong understanding of modern observability practices including logs, metrics, traces, monitoring, dashboards, and alerting.
  • Demonstrated experience with incident response, technical escalation, root-cause analysis, and problem management.
  • Strong understanding of FinOps and cloud financial management principles, including optimization, allocation, forecasting, and cost accountability.
  • Ability to evaluate architecture from both technical and financial perspectives.
  • Strong communication and collaboration skills with the ability to influence engineers, architects, Product leaders, Finance, Security, and executive stakeholders.

EEO Statement
iCIMS is a place where everyone belongs. We celebrate diversity and are committed to creating an inclusive environment for all employees. Our approach helps us to build a winning team that represents a variety of backgrounds, perspectives, and abilities. So, regardless of how your diversity expresses itself, you can find a home here at iCIMS.
We are proud to be an equal opportunity and affirmative action employer. We prohibit discrimination and harassment of any kind based on race, color, religion, national origin, sex (including pregnancy), sexual orientation, gender identity, gender expression, age, veteran status, genetic information, disability, or other applicable legally protected characteristics. If you would like to request an accommodation due to a disability, please contact us at
Compensation and Benefits
We accept applications for this position on an ongoing basis until the position is filled. Applications will be reviewed as they are received, and qualified candidates may be contacted throughout the posting period.
The anticipated base salary range for this position is $150,000 - $170,000. In addition, the estimated on-target earnings ("OTE"), which includes base salary and commissions, is $170,000-$200,000.
Actual compensation will depend on various job-related factors, including but not limited to, location, experience, and job qualifications. This range aligns with our commitment to equitable and transparent compensation practices, as required by applicable law.
Competitive health and wellness benefits include medical, dental, vision, 401(k), dependent care, short term and long-term disability, life and AD&D insurance, bonding and parental leave, mindfulness resources, an open vacation policy, sick days, paid holidays, quiet hours each workday, and tuition reimbursement. Benefits and eligibility may vary by location, role, and tenure. Learn more here:

Similar jobs