Haystack
← Back to Jobs
Engineering
NI

Observability Engineer

NimbusAITech LLCPhoenix, AZ🇺🇸United StatesPosted Sep 17, 2026

Quick Overview

Seniority
Mid Senior
Work mode
On Site
Location
Phoenix, AZ, United States
Posted
19 hours ago

Job Description

Job Title: Observability Engineer

Location: Phoenix, AZ (Onsite)

Job Type: Contract

 

Role Overview

We are seeking an experienced Observability Engineer to design, implement, and maintain enterprise-grade monitoring, logging, and distributed tracing solutions. In this onsite role based in Phoenix, AZ, you will partner closely with DevOps, SRE, and software engineering teams to ensure system reliability, improve incident response workflows, and optimize application performance across complex infrastructure environments.

 

Key Responsibilities

  • Telemetry & Monitoring Strategy: Architect and deploy end-to-end telemetry solutions covering metrics, logs, traces, and alert mechanisms across distributed systems.

 

  • Tooling Implementation: Configure, maintain, and scale observability platforms, primarily Dynatrace, Splunk, and OpenSearch / Elasticsearch.

 

  • Distributed Tracing: Implement OpenTelemetry (OTel) standards and instrumentation to achieve deep, service-level visibility across microservices and distributed applications.

 

  • Performance Optimization: Analyze application and infrastructure performance bottlenecks, establish baselines, and design proactive alerting thresholds to minimize downtime and MTTR.

 

  • SRE & Production Support: Collaborate with Site Reliability Engineering (SRE) and DevOps teams on incident management, blameless post-mortems, and root cause analysis (RCA).

 

  • Infrastructure as Code & Automation: Automate agent deployments, dashboard provisioning, alert rules, and remediation tasks using Python, Bash, Terraform, and Ansible.

 

Required Qualifications

  • 4+ years of hands-on experience in observability, SRE, systems engineering, or DevOps roles.
  • Deep technical expertise in enterprise monitoring platforms such as Dynatrace, Splunk, and Elasticsearch / OpenSearch.
  • Solid experience instrumenting applications with OpenTelemetry (OTel), distributed tracing, and custom metric collection.
  • Strong proficiency in scripting and automation with Python and Bash.
  • Hands-on infrastructure automation experience using Terraform and Ansible.
  • Proven track record supporting high-availability production systems and participating in structured incident response workflows.
  • Strong communication and analytical problem-solving skills; ability to work cross-functionally onsite.

 

Preferred Qualifications

 

  • Working knowledge of cloud infrastructure environments across AWS, Azure, or Google Cloud Platform.
  • Experience with containerization and orchestration platforms (Docker, Kubernetes).
  • Familiarity with SLO/SLI definition frameworks and error-budget tracking.

 

Similar jobs