Haystack
← Back to Jobs
Technology
RE

Staff Observability Platform Engineer (AI / GPU Infrastructure)

Recruitment.aiNew York, NY🇺🇸United StatesPosted 20 Aug 2026

Quick Overview

Seniority
Mid Senior
Work mode
On Site
Location
New York, NY, United States
Posted
Yesterday
KubernetesPrometheusPython

Job Description

Job Title: Staff Observability Platform Engineer (AI / GPU Infrastructure)

 Location: Seattle, WA / Houston, TX / New York, NY (Onsite)
 Employment Type: Contract


Position Overview

We are looking for a highly experienced Staff Observability Platform Engineer to design, build, and operate large-scale observability platforms supporting AI/GPU infrastructure and Kubernetes environments.

This is not a traditional monitoring or dashboard role. You''ll own the observability backend platform, ensuring it remains scalable, reliable, and cost-efficient as telemetry volumes grow.


Key Responsibilities

  • Design, build, and scale enterprise metricsloggingtracing, and telemetry platforms.
  • Architect and operate distributed observability backends using PrometheusMimirThanosVictoriaMetricsCortexLokiElasticsearch, or similar technologies.
  • Build and optimize OpenTelemetry Collector pipelines including routing, filtering, sampling, and exporters.
  • Optimize cardinalityingestionretentionstoragequery performance, and infrastructure cost.
  • Manage large-scale Kubernetes observability across multi-cluster environments.
  • Troubleshoot production issues involving metricslogstraces, networking, storage, and distributed systems.
  • Write and review production-quality code and establish observability standards across engineering teams.
  • Partner closely with Platform, Infrastructure, Security, and Application Engineering teams.

Required Qualifications

  • 8+ years of experience in ObservabilityPlatform EngineeringSRE, or Infrastructure Engineering.
  • Hands-on experience operating production-scale observability platforms such as MimirThanosVictoriaMetricsCortexLoki, or Elasticsearch.
  • Strong expertise with PrometheusOpenTelemetry, and telemetry pipeline design.
  • Experience managing large-scale production environments with measurable metrics (ingestion rates, active time series, storage, retention, cluster size, etc.).
  • Strong understanding of cardinalityretention strategiesstorage architecturesamplingquery optimization, and cost management.
  • Deep hands-on experience with Kubernetes, distributed systems, networking, service discovery, autoscaling, and reliability engineering.
  • Strong programming skills in Go and/orn Python.
  • Solid understanding of production engineering concepts including retries, backpressure, buffering, circuit breaking, graceful degradation, scalability, and failure handling.

Preferred Qualifications

  • Experience with GPU infrastructureAI/ML platforms, or HPC environments.
  • Knowledge of NVIDIA DCGMInfiniBandRoCE/RDMANVLinkNVSwitchNCCL, or Slurm.
  • Experience building custom Prometheus ExportersOpenTelemetry Collectors, or large multi-cluster Kubernetes observability platforms.

Ideal Candidate

We''re looking for someone who has personally owned and operated observability backends at production scale, understands the trade-offs between cardinality, retention, storage, performance, and cost, and can build scalable observability platforms for modern AI/GPU infrastructure. Experience limited to dashboards, alerts, or consuming monitoring tools without backend ownership will not be sufficient for this role.

Similar jobs