Haystack
← Back to Jobs
Technology
ST

Site Reliability Engineer II

SRI Tech SolutionsUnited States🇺🇸United StatesPosted 4 Sept 2026

Why This Role Stands Out

This Site Reliability Engineer II role offers significant growth potential by allowing you to work with cutting-edge Google Cloud Platform data platforms and develop advanced automation skills. You'll thrive if you're passionate about ensuring system reliability and proactively solving complex technical challenges in a collaborative, remote-friendly environment. Apply now to contribute to a reputable tech company and advance your SRE career.

Quick Overview

Seniority
Mid Senior
Work mode
Hybrid
Location
United States
Posted
21 hours ago
New RelicBigQueryDatadogGoogle CloudPython

Job Description

Site Reliability Engineer II

REMOTE 

Full Time

Overview / Summary

We are seeking a Site Reliability Engineer to join an SRE team focused on observability, monitoring, and technical consulting across Google Cloud Platform-based data platforms. This role is responsible for ensuring the availability, reliability, and performance of cloud and network systems and services through automation, monitoring, troubleshooting, and continuous optimization.

Key Responsibilities

  • Collaborate with infrastructure teams to implement critical solutions by automating routine tasks.
  • Monitor and manage production environments, proactively identifying and resolving issues.
  • Participate in building advanced tooling for system access monitoring, log session recording, and reliability administration across multiple geographically distributed data centers.
  • Engage with engineering teams to improve on-call efficiencies, incident management, and post-mortem analysis.
  • Perform capacity planning and optimization to support growing demands and traffic patterns.
  • Maintain monitoring and alerting systems for proactive system health checks.
  • Continuously improve system performance, stability, and security through data-driven analysis and optimization.
  • Create and maintain comprehensive documentation and diagrams to facilitate knowledge sharing.
  • Work hands-on with cloud infrastructure, BigQuery workloads, CI/CD pipelines, and enterprise monitoring tools to maintain critical systems at scale.

Required Qualifications

  • Bachelor's degree.
  • 4+ years of experience in IT.
  • 3+ years of development experience.
  • Practitioner-level experience with at least one coding language or framework.
  • Hands-on experience with Google Cloud Platform (Google Cloud Platform).
  • Experience with BigQuery.
  • Experience with Dynatrace.
  • Proficiency with monitoring and observability tools, ideally Dynatrace or comparable tools such as Datadog or New Relic.
  • Familiarity with ITSM tools such as ServiceNow, including incident, problem, and change management.

Preferred Qualifications

  • Experience with Google Cloud Platform Cloud Run.
  • Experience with Python.
  • Strong troubleshooting and problem-solving skills.
  • Familiarity with AI tools, including agents, skills, LLMs, and copilots.
  • Experience defining and tracking SLAs, SLOs, and SLIs.

Similar jobs