Observability Architect - Seattle, Alpharetta or Cincinnati
Quick Overview
Job Description
DTS is looking for Observability Architect for our Client position based in Seattle, Alpharetta or Cincinnati
Job Description
Overview:
Observability & Enterprise Monitoring Architect with specialized expertise in SolarWinds platform architecture, design, and broader multi-tool observability ecosystems. Working knowledge of OpenText NNMi will be an added advantage. This role will be responsible for the end-to-end architecture, deployment, implementation, optimization, integration, and operational governance of enterprise-scale implementations of monitoring solutions (SolarWinds). Responsible for deploying platform infrastructure, establishing platform health standards, architecting automated alert workflows, designing hybrid/cloud monitoring integrations, and collaborating closely with cross-functional infrastructure and leadership teams to ensure high availability, scalability, and performance.
Roles & Responsibilities:
- Platform Architecture, Deployment & Lifecycle Management (SolarWinds)
- Core Module Architecture & Deployment: Design, deploy, configure, and optimize SolarWinds modules including NPM, NCM, NTA, SAM, and the broader Orion / SWOSH (Hybrid Cloud Observability) platform ecosystem.
- Deployment & Upgrade Strategy: Lead new platform rollouts, migrations, and routine/major version updates across platform components; establish standards for platform health governance using Active Diagnostics and My Deployment health checks.
- Polling Infrastructure Deployment: Architect, deploy, scale, and load-balance Additional Polling Engines (APEs) to ensure optimal performance, redundancy, and capacity across enterprise environments.
- Database & Storage Strategy: Oversee architectural strategy for the underlying MS SQL Database, ensuring high availability, performance tuning, and robust configuration and database backup governance.
- Network & Device Monitoring
- Discovery & Asset Onboarding: Execute network discoveries, deploy automated node onboarding/offboarding frameworks, assign Universal Device Pollers (UnDP), and maintain custom attribute taxonomies and group hierarchies.
- Configuration Governance (NCM): Design and implement NCM command templates, establish policies for automated daily startup/running config backups, config archiving, and remediation frameworks for compliance/transfer failures.
- Topology & Visualization: Build dynamic, accurate network topology frameworks using Network Atlas and modern visual canvases aligned with enterprise requirements.
- Alert Architecture, Dashboarding & ITSM Integration
- Signal & Alert Optimization: Design, implement, and tune custom Alert Triggers, Actions, and Threshold frameworks to eliminate alert noise and establish high-signal, actionable alerting.
- ITSM & Workflow Deployment: Deploy bi-directional ITSM/ticketing integrations to enable automated ticket creation, enrichment, routing, and lifecycle tracking.
- Reporting & Visibility Frameworks: Build enterprise operational and executive Dashboards, Views, and Reports tailored to multi-level stakeholder requirements.
- Incident & Deployment Support: Lead technical reviews for complex operational anomalies, troubleshoot systemic telemetry or deployment issues, and collaborate with domain teams on root cause analysis (RCA).
- AIOps & Next-Gen Operations
- AIOps Implementation: Define, deploy, and leverage AIOps, machine learning, and pattern-recognition capabilities to identify baseline anomalies, reduce event noise, and drive predictive incident management.
- Automated Remediation Architecture: Collaborate with cross-functional teams to integrate AI-driven event correlation models and deploy automated self-healing remediation workflows into the central monitoring platform.
- Integration, Vendor Coordination
- Integration, Vendor Coordination
- API & Integration Deployment: Implement REST API and webhook integration models across enterprise applications, tools, and platforms as per business requirements.
- Troubleshooting & Diagnostics
- Advanced Escalation: Perform deep-dive troubleshooting and root-cause analysis for complex, platform-level performance degradations, engine polling deadlocks, and monitoring agent corruptions.
- Telemetry Diagnostics: Utilize Active Diagnostics and system telemetry data to investigate and resolve complex network configuration transfer failures, polling sync latency, and data ingestion bottlenecks.
Required Skills:
- Multi-tool Architecture & Deployment Expertise: Deep architectural and hands-on implementation knowledge of enterprise monitoring tools (SolarWinds, OpenText, Splunk, etc.) at global scale.
- Protocol & Telemetry Mastery: Advanced understanding of SNMP (v2c/v3), WMI, WinRM, Syslog, NetFlow/sFlow, and core Observability pillars (Metrics, Logs, Traces).
- Automation & API Design: Intermediate skills in PowerShell/Python, REST APIs, and building API-driven automation for enterprise monitoring workflows.
- AIOps & Intelligent Automation: Strong grasp of AIOps concepts, machine learning algorithms for anomaly detection, automated event correlation, and predictive analytics within modern observability frameworks.
- Cloud & Hybrid Deployment: Hands-on experience architecting and deploying enterprise platform monitoring into AWS, Azure, or Google Cloud Platform environments.
- Infrastructure Foundations:
- System Administration: Advanced knowledge of Windows and Linux platform architecture and administration.
- Database Architecture: In-depth understanding of MS SQL/Database architecture, performance tuning, and query execution.
- Networking: Comprehensive understanding of enterprise networking architectures including TCP/IP, DNS, DHCP, Routing, and Switching.
- ITSM: Deep experience in enterprise ITSM processes, ITIL frameworks, and operational governance.
DTS offers excellent compensation package.
Contact:
Karun Sharma
Team Lead
Digital Technology Solutions (DTS)
Skills
Similar jobs
LabVIEW Programmer
Mission Microwave · Cypress, United States
8 minutes agoPhotography Intern
JOHN MORAN AUCTIONEERS INC · Monrovia, United States
8 minutes agoDirector (Healthcare)
Abacus Service Corporation · San Francisco, United States
3 hours agoChange Management Specialist (ServiceNow)
Kriscon · Buffalo, United States
3 hours agoUser Acceptance Testing Analyst
Link Technologies · Las Vegas, United States
3 hours agoDatabricks Architect
Theron Partners Inc. · Dallas, United States
3 hours ago