Why This Role Stands Out
This hybrid Sr. Site Reliability Engineer role offers a competitive salary and the chance to significantly impact a leading financial software company by enhancing system reliability and observability. You'll thrive here if you are passionate about infrastructure as code, performance monitoring with New Relic, and collaborating with engineering teams to build robust, scalable solutions. Don't miss this opportunity to advance your career in a dynamic and supportive environment.
Quick Overview
Job Description
Location: Wheeling, IL Salary: $150,000.00 USD Annually - $185,000.00 USD Annually Description:
Financial Software company
Position: SRE Engineer
Location: Hybrid remote/Buffalo Grove/Chicago
Comp: Solid Base, bonus, equity
Benefits: Comprehensive Health, 401K and Equity plan
Design and implement monitoring, alerting, and dashboards in New Relic (APM, Infrastructure, Logs, Synthetics) across Azure and AWS; write NRQL queries for troubleshooting, analysis, and reporting.
Define and implement SLOs/SLIs and error budgets; coach teams on using them to balance feature velocity with reliability and communicate system health to stakeholders.
Lead alert noise reduction and signal quality engineering-tune thresholds, eliminate false positives, and ensure every alert is actionable.
Optimize observability costs through log ingestion management, pipeline rules, and New Relic configuration governance.
Partner with engineering teams to improve observability maturity: structured logging, metrics instrumentation (RED/USE methods), distributed tracing, and effective dashboard patterns.
Develop and maintain Terraform infrastructure as code for provisioning and managing monitoring resources, alert configurations, and observability infrastructure-this is a primary engineering responsibility, not an occasional task.
Establish and enforce IaC governance standards for observability infrastructure across teams, providing a repeatable, auditable model for how monitoring resources are managed.
Author and troubleshoot Azure DevOps pipelines; support teams with deployment visibility, change tracking, and release hygiene as it relates to production reliability.
Administer and configure Incident.
IO: alert routing, notification workflows, Slack and OpsGenie integration, and runbook management-operationalizing what exists today and expanding from there.
Build out incident management foundations that are largely yours to establish: PIR/postmortem processes, on-call rotation design, escalation policies, incident severity classification, and response playbooks.
Track and report on MTTR, MTTD, and incident frequency; identify trends and drive continuous improvement in partnership with engineering teams.
Respond to and debrief on production incidents-providing real-time troubleshooting support and facilitating structured post-incident reviews.
Enable stream-aligned engineering teams to adopt improved observability and incident management practices through workshops, consultation, and hands-on guidance.
Collaborate with the Subsystems Platform Team to translate common needs into self-service observability and incident management capabilities.
Build lasting team competency through documentation, training materials, and knowledge-sharing sessions that outlast any individual engagement.
WHAT YOU'LL BRING
Core SRE Experience
By providing your phone number, you consent to: (1) receive automated text messages and calls from the Judge Group, Inc. and its affiliates (collectively "Judge") to such phone number regarding job opportunities, your job application, and for other related purposes. Message & data rates apply and message frequency may vary. Consistent with Judge's Privacy Policy, information obtained from your consent will not be shared with third parties for marketing/promotional purposes. Reply STOP to opt out of receiving telephone calls and text messages from Judge and HELP for help.
Contact
This job and many more are available through The Judge Group. Please apply with us today!
Similar jobs
- RA
Platform Engineer II
NewRaytheon
Marlborough, MA🇺🇸HybridYesterdaySplunkAgileBash+11Technology - CG
DevOps Software Engineer with AWS with Security Clearance
NewCCS Global Tech
Bethesda, MD🇺🇸$100k - $180k/yrOn-siteYesterdayMicroservicesAWSELK+15Technology - I3
Hybrid Cloud Platform Engineer with Security Clearance
Newi3
Huntsville, AL🇺🇸HybridYesterdayAWSEncryptionAnsible+8Technology - NT
Principal, SRE
NewNorthern Trust
Chicago, IL🇺🇸$137.4k - $233.6k/yrHybridYesterdayAnsibleC#Chef+4Technology - JM
Lead Platform SRE
NewJ.P. Morgan
Jersey City, New Jersey🇺🇸On-site1 hour agoSpringSpring BootAWS+3Technology - PG
AutoSys Platform Engineer
NewPTR Global
Irvine, CA🇺🇸$65 - $70/hrOn-siteYesterdayShellAWSKubernetesTechnology