Haystack
← Back to Jobs
Technology

Senior Couchbase DevOps & SRE Engineer // Sunnyvale, CA

Alliance ITSunnyvale, CA🇺🇸United StatesPosted 18 Aug 2026

Why This Role Stands Out

This hybrid role offers a fantastic opportunity to design, automate, and scale critical Couchbase database infrastructure, allowing you to significantly impact a reputable company's technology landscape. If you thrive on SRE principles, possess strong Couchbase and DevOps expertise, and enjoy a collaborative environment, this position is an excellent next step for your career growth. Embrace this chance to innovate and advance your skills in a dynamic setting.

Quick Overview

Work Type
Hybrid
Level
Mid Senior

Job Description

We’re Hiring: Senior Couchbase DevOps & SRE Engineer

We are looking for a Senior Couchbase DevOps & SRE Engineer to design, operate, automate, and scale highly available Couchbase database infrastructure.

The ideal candidate will have strong experience in Couchbase administration, DevOps, SRE practices, automation, infrastructure as code, monitoring, and production operations.

🔹 Key Responsibilities

Infrastructure & Couchbase Operations

  • Design, implement, and maintain highly available, scalable Couchbase clusters across multi-node and multi-zone environments.
  • Manage cluster operations including rebalancing, failover, scaling, patching, and XDCR configuration.
  • Perform capacity planning, bucket sizing, memory optimization, and storage management.

CI/CD & Infrastructure as Code

  • Develop and maintain CI/CD pipelines for application and Couchbase infrastructure deployments.
  • Implement Infrastructure as Code using Terraform or CloudFormation for cluster provisioning and configuration management.
  • Automate rolling upgrades, patching, and zero-downtime deployments.

Automation & SRE

  • Define and implement SRE practices including SLIs, SLOs, error budgets, and reliability objectives.
  • Identify operational bottlenecks and build automation to eliminate manual Couchbase and platform tasks.
  • Automate backup, restore, XDCR setup, and cluster health validation.

Monitoring & Incident Management

  • Implement proactive monitoring and alerting for cluster health, N1QL performance, replication lag, and memory pressure using Prometheus, Grafana, or Datadog.
  • Lead incident response, root-cause analysis (RCA), and preventive actions for Couchbase production issues.
  • Develop and maintain operational runbooks and SOPs.

Security & Compliance

  • Manage Couchbase RBAC, TLS/SSL, audit logging, and encryption at rest.
  • Enforce security hardening standards across all Couchbase environments.

Skills

Encryption
CloudFormation
Datadog
Grafana
Prometheus
REST
Terraform

Similar jobs