Haystack
← Back to Jobs
Remote
Engineering
IC

Senior Observability & Monitoring Engineer

Infinite Computer Solutions (ICS)United States🇺🇸United StatesPosted 1 Sept 2026

Why This Role Stands Out

This remote Senior Observability & Monitoring Engineer role offers a fantastic opportunity to shape critical cloud migration initiatives and enhance system reliability using cutting-edge tools like Dynatrace and Splunk. You'll thrive here if you're an experienced engineer passionate about building robust monitoring solutions and collaborating with diverse teams to achieve operational excellence. Apply today to make a significant impact on mission-critical applications!

Quick Overview

Seniority
Mid Senior
Work mode
Remote
Location
United States
Posted
1 week ago
Jira

Job Description

Senior Observability & Monitoring Engineer
Remote Role
Fulltime/ W2- Infinite Computer Solutions
About the Role
We are seeking an experienced Senior Observability & Monitoring Engineer to design, implement, and optimize enterprise monitoring and observability solutions supporting the migration of mission-critical applications from on-premises environments to AWS and Google Cloud Platform (Google Cloud Platform).
You will play a key role in improving system reliability by enabling proactive monitoring, intelligent alerting, rapid incident detection, and operational excellence.
Key Responsibilities
  • Design and implement enterprise-wide monitoring and observability solutions.
  • Build dashboards, KPIs, SLIs, and SLOs for applications, infrastructure, databases, APIs, and cloud services.
  • Configure intelligent monitoring, alerting, event correlation, and anomaly detection using Dynatrace, Splunk, and Moogsoft.
  • Develop synthetic monitoring and automated health checks for critical business services.
  • Partner with Cloud Architects, Developers, SREs, and Operations teams to ensure production readiness during cloud migrations.
  • Automate monitoring configurations and operational processes using scripting and AI-assisted tools.
Required Qualifications
  • 8+ years of experience in Monitoring, Observability, SRE, or Operations Engineering.
  • Strong hands-on experience with Dynatrace, Splunk, and Moogsoft.
  • Experience supporting cloud migrations to AWS and/or Google Cloud Platform.
  • Solid understanding of APM, distributed tracing, logging, metrics, and alert management.
  • Experience with Jira, ServiceNow, and incident management.
  • Scripting skills in Python, Bash, or PowerShell.
  • Strong knowledge of enterprise applications and distributed systems.
Preferred Qualifications
  • Experience in fintech, banking, or other regulated industries.
  • Knowledge of OpenTelemetry, Terraform, and Infrastructure as Code (IaC).
  • Experience with self-healing automation and cloud-native observability.
What Success Looks Like
  • Establish scalable monitoring standards for cloud migration initiatives.
  • Improve issue detection and reduce Mean Time to Resolution (MTTR).
  • Minimize alert fatigue through intelligent alert optimization.
  • Deliver reusable observability frameworks that enhance operational reliability.

Similar jobs