Haystack
← Back to Jobs
Technology

Senior Couchbase DevOps & SRE Engineer // Sunnyvale, CA and Austin, TX

Alliance ITManor, TX🇺🇸United StatesPosted 18 Aug 2026

Why This Role Stands Out

This role offers a fantastic opportunity to shape and scale critical database infrastructure, providing significant career growth in DevOps and SRE practices. You'll thrive here if you have a strong background in Couchbase administration and a passion for automation and highly available systems, all within a hybrid work environment. Apply today to join a dynamic team and make a tangible impact!

Quick Overview

Work Type
Hybrid
Level
Mid Senior

Job Description

We’re Hiring: Senior Couchbase DevOps & SRE Engineer

We are looking for a Senior Couchbase DevOps & SRE Engineer to design, operate, automate, and scale highly available Couchbase database infrastructure.

The ideal candidate will have strong experience in Couchbase administration, DevOps, SRE practices, automation, infrastructure as code, monitoring, and production operations.

🔹 Key Responsibilities

Infrastructure & Couchbase Operations

  • Design, implement, and maintain highly available, scalable Couchbase clusters across multi-node and multi-zone environments.
  • Manage cluster operations including rebalancing, failover, scaling, patching, and XDCR configuration.
  • Perform capacity planning, bucket sizing, memory optimization, and storage management.

CI/CD & Infrastructure as Code

  • Develop and maintain CI/CD pipelines for application and Couchbase infrastructure deployments.
  • Implement Infrastructure as Code using Terraform or CloudFormation for cluster provisioning and configuration management.
  • Automate rolling upgrades, patching, and zero-downtime deployments.

Automation & SRE

  • Define and implement SRE practices including SLIs, SLOs, error budgets, and reliability objectives.
  • Identify operational bottlenecks and build automation to eliminate manual Couchbase and platform tasks.
  • Automate backup, restore, XDCR setup, and cluster health validation.

Monitoring & Incident Management

  • Implement proactive monitoring and alerting for cluster health, N1QL performance, replication lag, and memory pressure using Prometheus, Grafana, or Datadog.
  • Lead incident response, root-cause analysis (RCA), and preventive actions for Couchbase production issues.
  • Develop and maintain operational runbooks and SOPs.

Security & Compliance

  • Manage Couchbase RBAC, TLS/SSL, audit logging, and encryption at rest.
  • Enforce security hardening standards across all Couchbase environments.

Skills

Encryption
CloudFormation
Datadog
Grafana
Prometheus
REST
Terraform

Similar jobs