Senior Manager – Observability Engineering- 10+ yrs- Atlanta, United States- Onsite
Why This Role Stands Out
This role offers a fantastic opportunity to lead and shape the observability strategy for enterprise platforms, leveraging cutting-edge AI technologies to drive operational excellence. You'll thrive here if you're an experienced engineering leader passionate about building and mentoring high-performing teams, with a proven track record in observability and cloud environments. Apply now to make a significant impact and advance your career in a dynamic, onsite position.
Quick Overview
Job Description
Job Requirements Key Responsibilities Look for Local candidates to Atlanta it is 5 days onsite role. Lead and mentor a team of observability engineers supporting enterprise platforms and services.Define and execute the observability strategy, standards, and roadmap.Oversee monitoring, logging, alerting, tracing, and dashboarding solutions.Drive service reliability, incident response readiness, and operational excellence initiatives.Collaborate with application, infrastructure, cloud, and SRE teams to improve system health and performance.Establish KPIs, SLAs, and operational metrics to measure platform reliability and team effectiveness.Manage hiring, performance development, resource planning, and stakeholder communications.Ensure adoption of best practices for observability, automation, and proactive problem management.Champion the adoption of AI and Copilot-enabled workflows within the Observability organization.Evaluate and implement AI-driven monitoring, alert correlation, and incident management capabilities.Partner with engineering and platform teams to build intelligent operational dashboards and automated remediation solutions.QualificationsBachelor's degree in Computer Science, Engineering, or a related field.10+ years of experience in infrastructure, operations, SRE, platform engineering, or observability domains.3+ years of people management experience leading technical teams.Strong knowledge of observability platforms such as Splunk, Datadog, AppDynamics, Dynatrace, Grafana, Prometheus, OpenTelemetry, or similar tools.Experience working in cloud environments (Azure & Google Cloud Platform).Experience leveraging Microsoft Copilot, Generative AI, and AI-powered observability capabilities to improve operational efficiency, incident response, and engineering productivity.Knowledge of AI-assisted troubleshooting, anomaly detection, root cause analysis, and predictive monitoring solutions.Excellent communication, stakeholder management, and leadership skills.PreferredExperience leading globally distributed teams.Strong background in automation, DevOps, and reliability engineering practices.Familiarity with enterprise-scale monitoring and incident management processes.This role will be responsible for building a high-performing observability team that enables proactive detection, rapid troubleshooting, and improved reliability across business-critical services.Work Experience 10-15Years
Similar jobs
- VS
Construction Project Engineer
NewVarda Space Industries
El Segundo🇺🇸14 hours agoFiberArticulateAutoCAD+6Engineering - ST
Engineering Manager - FinTech
NewStubHub
Los Angeles🇺🇸Yesterday401kSwiftCompliance+6Engineering - SE
Intern, Engineering Sciences (Summer 2027)
NewSecretariat
Denver🇺🇸3 hours agoHTTPSMicrosoft OfficeEngineering - SE
Engineering Manager, SupportX
SeatGeek
New York🇺🇸1 week agoDynamoDBFastAPIMemcached+17Engineering - KR
PROJECT ENGINEER
NewKnife River
Long Beach, CA🇺🇸On-site19 hours agoAutoCADMicrosoft OfficeEngineering - SW
Plant Maintenance Engineer (Equipment)
NewSmurfit Westrock
Jordan, NY🇺🇸$39 - $43/hrHybrid19 hours agoEngineering