Haystack
← Back to Jobs
Remote
Technology

Site Reliability Engineer – OpenSearch - Remote

The Dignify Solutions, LLCUnited States🇺🇸United StatesPosted 6 Aug 2026

Why This Role Stands Out

This remote Site Reliability Engineer role offers an exceptional opportunity to shape the reliability and performance of critical cloud services for a reputable company. You'll thrive here if you possess deep expertise in OpenSearch, Kubernetes, and automation, and are eager to contribute to a dynamic team focused on innovation and excellence.

Quick Overview

Work Type
Remote
Level
Mid Senior

Job Description

  • 10+ years of experience in Site Reliability Engineer – OpenSearch to help ensure the highest levels of availability, performance, scalability, and Quality of Service (QoS) for mission-critical cloud services.
  • Deep experience in site reliability engineering, DevOps, cloud operations, automation, observability, and distributed systems, with proven hands-on expertise architecting, building, deploying, operating, and optimizing high-performance OpenSearch clusters and platforms from the ground up in production environments.
  • Expert with Kubernetes, including troubleshooting, operations, management, and configuration of complex Kubernetes services.
  • Proven hands-on expertise designing, building, deploying, supporting, and maintaining OpenSearch clusters and platforms from scratch in production environments
  • Strong experience with OpenSearch administration, cluster architecture, performance tuning, scaling, upgrades, and troubleshooting
  • Experience with index design, shard and replica strategy, cluster sizing, node management, snapshot/restore, backup, and disaster recovery
  • Strong understanding of distributed systems, search platforms, indexing pipelines, query optimization, and high-availability architectures
  • Expertise with Git
  • Expertise with Concourse, including setup, management, and troubleshooting of new pipelines
  • Expertise with Linux, specifically SUSE and Ubuntu
  • Expertise with Kafka, Zookeeper, and Big Data technologies
  • Expert in development of automation for testing, deployment, scalability, and management of cloud services
  • Expertise with building, implementing, and/or supporting cloud monitoring tools
  • Expert knowledge of cloud computing, infrastructure operations, and databases
  • Expert understanding of web services, networking, virtualization, and internet protocols
  • Ability to multitask and handle various projects, deadlines, and changing priorities
  • Excellent communication and prioritization skills
  • Expertise with security fundamentals as they pertain to SaaS multi-tenant application systems
  • Experience with AWS services including Route 53, EC2, S3, CloudWatch, DynamoDB, RDS, IAM, ACM, KMS, and VPC
  • Experience deploying and operating OpenSearch in AWS-based environments
  • Experience with Cloud Foundry-based environments
  • Experience with Jenkins, Chef, and/or Terraform
  • Exposure to and understanding of troubleshooting IP networks and application stacks
  • Experience with observability tools such as Prometheus and Grafana
  • Experience with log ingestion pipelines, index lifecycle management, retention strategies, and search platform security controls
 

 

Skills

DynamoDB
AWS
Chef
Git
Grafana
Jenkins
Kafka
Kubernetes
Prometheus
Terraform

Similar jobs