Haystack
← Back to Jobs
Engineering
ST

Grafana & Observability Engineer - Dallas, Tampa & Jersey City

StradITDallas, Texas🇺🇸United StatesPosted 20 Aug 2026

Quick Overview

Seniority
Mid Senior
Work mode
Hybrid
Location
Dallas, Texas, United States
Posted
10 hours ago
AWSNew RelicSplunkBashContinuous ImprovementGrafanaKubernetesOnboardingPowerShellPrometheusPythonRoot Cause AnalysisTerraform

Job Description

Role: Grafana & Observability Engineer

Experience: 5 to 10 years

Employment: W2

Location: Jersey City NJ, Tampa FL & Dallas TX

Key Responsibilities

Observability Platform Engineering

  • Administer and support Grafana Cloud and on-premises Grafana deployments.
  • Design and implement enterprise observability solutions for metrics, logs, traces, synthetic monitoring, and alerting.
  • Establish and maintain observability standards, best practices, and governance processes.
  • Configure and manage Grafana data sources, alerting, RBAC, folders, teams, and integrations.
  • Ensure platform scalability, reliability, resiliency, and operational excellence.

Automation & Infrastructure as Code

  • Develop and maintain Terraform modules for Grafana infrastructure and configuration management.
  • Automate onboarding of applications, infrastructure, dashboards, alerts, and data sources.
  • Build self-service capabilities that reduce manual operational effort and improve adoption.
  • Integrate observability capabilities into CI/CD and infrastructure provisioning workflows.

Monitoring, Alerting & Incident Management

  • Design meaningful monitoring and alerting strategies based on service health and business-critical workflows.
  • Implement and optimize alerting standards to reduce noise and improve signal quality.
  • Support incident response, troubleshooting, root cause analysis, and post-incident reviews.
  • Drive continuous improvement of operational visibility and platform health.

OpenTelemetry & Telemetry Engineering

  • Implement and support OpenTelemetry instrumentation across applications and infrastructure.
  • Establish standards for logs, metrics, traces, and telemetry collection.
  • Support telemetry pipelines, agent deployments, and data collection strategies.
  • Assist teams with instrumentation design and observability adoption.

Migration & Modernization

  • Support migration initiatives from legacy observability platforms to Grafana.
  • Analyze existing monitoring, alerting, logging, and tracing implementations and recommend modernization approaches.
  • Develop reusable migration patterns, automation, and engineering standards.
  • Partner with application teams to accelerate adoption of enterprise observability capabilities.

Collaboration & Leadership

  • Work closely with application development, infrastructure, cloud, and SRE teams.
  • Provide technical leadership and mentoring to engineers across the organization.
  • Contribute to observability architecture, strategy, and roadmap development.
  • Promote observability as a core engineering practice across the enterprise.

Required Qualifications

  • Bachelor's degree in Computer Science, Engineering, Information Systems, or related field.
  • 5+ years of experience in observability, monitoring, operations, or platform engineering.
  • Hands-on experience administering Grafana in large-scale enterprise environments.
  • Strong experience with Terraform and Infrastructure as Code practices.
  • Experience implementing monitoring, alerting, logging, and distributed tracing solutions.
  • Experience with OpenTelemetry concepts, instrumentation, and telemetry pipelines.
  • Strong Linux and cloud platform administration skills.
  • Experience with scripting and automation using Python, PowerShell, Bash, or similar languages.
  • Knowledge of operational excellence, reliability engineering, and incident management practices.

Preferred Qualifications

  • Experience migrating from tools such as Splunk, Dynatrace, AppDynamics, New Relic, OpenText OBM, or similar platforms.
  • Experience with Grafana Alloy, Tempo, Loki, Mimir, or Prometheus.
  • Experience operating observability platforms in AWS environments.
  • Knowledge of Kubernetes, containers, and cloud-native observability.
  • Experience designing enterprise observability strategies and governance models.
  • Familiarity with CI/CD platforms and DevOps practices.

Desired Skills

  • Grafana Administration
  • Terraform
  • OpenTelemetry (OTEL)
  • Monitoring & Alerting
  • Observability Engineering
  • Platform Engineering
  • Linux Administration
  • AWS Cloud Services
  • Automation & Scripting
  • Incident Management
  • Infrastructure as Code
  • Telemetry Pipelines
  • Reliability Engineering
  • Root Cause Analysis
  • Enterprise Monitoring Architecture

Similar jobs