Haystack
← Back to Jobs
Technology
NM

Site Reliability Engineer-Prometheus, Grafana, and OpenTelemetry (W2 Only)

New Millennium ConsultingUnited States🇺🇸United StatesPosted Sep 25, 2026

Quick Overview

Salary
$50/hr
Seniority
Mid Senior
Work mode
Hybrid
Location
United States
Posted
Yesterday
AgileGrafanaJavaPrometheus

Job Description

One of my clients in Texas, is urgently looking for Site Reliability Engineer.

Hourly: $50 per hour (W2)

Duration: 6-8 months + extension

Location: Remote

Required Skills:

  1. 4-6 years of relevant work experience
  2. Experience with private, public and hybrid cloud environments
  3. Hands-on experience establishing Observability using Prometheus, Grafana, and OpenTelemetry
  4. Proven software development experience with a focus on building internal tooling and applying AIOps to streamline and accelerate development
  5. Experience with object-oriented programming (preferably Java), cloud architecture, CI/CD pipelines, and modern design patterns
  6. Experience with FinOps practices, including cloud cost allocation and resource optimization
  7. Experience with security frameworks for user and services authorization and authentication
  8. Experience with destructive and performance test design and execution
  9. Experience with modern debugging and root cause analysis techniques
  10. Experience with version control and code repositories, such as GitHub

Scope:

50% Delivery and Execution

  • Builds, scales, and maintains robust CI/CD pipelines, workflows, and cloud infrastructure, with a clear understanding of Reliability Engineering practice areas, to ensure deployments meet all resiliency, security and change control requirements. Takes on new opportunities and tough challenges with a sense of urgency, high energy and enthusiasm. Consistently achieves results, even under tough circumstances
  • Implements Infrastructure as Code (IaC) and configuration management to ensure scalable, immutable, and reproducible build and deployment environments across product lifecycles
  • Drives automation and toil reduction across the build ecosystem, provisioning, artifact management, and deployment pipelines
  • Implements SLOs/SLIs for Critical User Journeys (CUJs) and comprehensive observability(metrics, alerting, and distributed tracing) for cloud infrastructure and applications

20% Learns and Grows

  • Learns through successful and failed experiment when tackling new problems. Actively seeks ways to grow and be challenged using both formal and informal development channels
  • Continuously evaluates emerging AIOps, DevOps, and cloud-native build technologies to enhance performance, security, and time to detection/resolution

20% Plans and Aligns

  • Collaborates with other team members in agile processes. Creates new and better ways for the organization to be successful. Works the Product Team to ensure user stories are valuable, developer ready, easy to understand and testable. Delivers multi-mode communications that convey a clear understanding of the unique needs of different audiences. Adapts approach and demeanor in real time to match the shifting demands of different situations. Relates openly and comfortably with diverse groups of people
  • Partners with application engineering and support teams to meet reliability standards, optimize build/release workflows, and eliminate developer bottlenecks
  • Represents RE team in incident response efforts for infrastructure outages and in blameless post-mortems (BPMs)to prevent recurrence

10% Supports and Enables

  • Participates in on-call rotation to support customer intake requests and serves as first point of contact for RE Build team during incident responses
  • Helps grow junior-level engineers and contractors by providing guidance on best practices, leading technical discussions and conducting knowledge transfers

Similar jobs