Quick Overview
Job Description
Position Summary
We are seeking an experienced Database Site Reliability Engineer (SRE) to support and operate mission-critical database platforms within a fast-paced enterprise environment. This role is focused on operational excellence, reliability, resiliency, automation, and continuous improvement across multiple database technologies.
This is not a development role. We are looking for hands-on database professionals who thrive in production operations, take ownership of issues, understand urgency, and are passionate about improving systems, processes, and themselves.
The ideal candidate views reliability as a product, proactively identifies risks before they become incidents, and continuously seeks opportunities to automate repetitive tasks and improve platform stability.
Key Responsibilities
Database Operations & Reliability
- Install, configure, upgrade, patch, and maintain enterprise database platforms.
- Ensure availability, performance, recoverability, and security of production database environments.
- Monitor database platforms and associated infrastructure, responding rapidly to incidents and service degradations.
- Lead troubleshooting efforts for database, operating system, storage, replication, and application connectivity issues.
- Execute failovers, disaster recovery testing, and recovery procedures.
- Partner with application teams to provide database guidance and operational support.
Platform Engineering
- Deploy, maintain, and optimize database infrastructure across physical, virtual, and cloud environments.
- Implement scalable, resilient database solutions.
- Evaluate and recommend improvements to architecture, monitoring, automation, and operational processes.
- Support capacity planning, performance tuning, and platform lifecycle management.
Automation & Continuous Improvement
- Develop and maintain automation solutions using Python, Shell, Ansible, or similar technologies.
- Help eliminate manual operational activities through engineering and automation.
- Improve monitoring, alerting, reporting, and operational workflows.
- Drive incremental improvements that reduce risk, improve reliability, and increase operational efficiency.
Performance & Incident Management
- Analyze and resolve database performance issues.
- Troubleshoot replication, backup/recovery, storage, network, and infrastructure-related incidents.
- Participate in root cause analysis and drive permanent corrective actions.
- Review operational metrics and trends to identify opportunities for improvement.
Operational Excellence
- Maintain accurate operational documentation, standards, and procedures.
- Generate and present operational metrics, service health indicators, and reliability reporting.
- Participate in incident response activities.
- Demonstrate strong ownership from issue identification through resolution.
Required Qualifications: -
- Strong experience administering enterprise database platforms, including:
- Sybase ASE
- Oracle RAC
- Additional database technologies such as MongoDB, Cassandra, Redis, PostgreSQL, MySQL, or similar platforms are a plus.
- Experience performing:
- Installation
- Configuration
- Upgrades
- Patching
- Performance tuning
- Backup and recovery
- High availability and disaster recovery
- Experience with database replication technologies including:
- SAP Replication Server
- Data Guard
- HVR (preferred)
- Strong Linux administration skills.
- Experience with automation and scripting:
- Python
- Ansible
- Shell scripting
- Understanding of storage, networking, operating systems, and infrastructure services.
- Experience with Veritas Cluster Server, ASM, LVM, and SAN technologies.
- Familiarity with enterprise operational tooling such as Jira, Service Now and Confluence.
- Strong analytical, troubleshooting, and problem-solving skills.
What Success Looks Like
The successful candidate:
- Takes ownership and drives issues to closure.
- Understands the urgency required to support critical production environments.
- Continuously improves systems, processes, and operational effectiveness.
- Learns quickly and adapts to new technologies.
- Balances operational stability with engineering innovation.
- Communicates clearly and effectively during incidents and high-pressure situations.
- Demonstrates a strong sense of accountability and professionalism.
- Leaves the platform better than they found it every day.
Preferred Mindset
We hire for attitude as much as technical skill.
We're looking for individuals who are:
- Customer obsessed
- Accountable and dependable
- Urgent without being reckless
- Continuously learning
- Driven to automate repetitive work
- Detail-oriented
- Collaborative but willing to lead
Focused on long-term platform reliability rather than short-term fixes
Similar jobs
- CA
GenAI Platform Engineer (Python, MCP & Agentic AI)
NewCapgemini America, Inc.
NY🇺🇸On-siteYesterdayFastAPIAWSGenerative AI+4Technology - VB
Network Engineer
Vaco by Highspring
Rollingwood, TX🇺🇸On-site3 weeks agoAWSAzurePowerShell+1Technology - IN
SRE Engineer
NewIntraedge
Phoenix, AZ🇺🇸On-siteYesterdayAPI GatewayService MeshBash+9Technology - MR
DevOps / SRE Engineer
NewMotion Recruitment Partners, LLC
Danvers, MA🇺🇸On-siteYesterdayBashJenkinsJira+1Technology - MA
OpenShift Engineer- W2
Magicforce
United States🇺🇸Hybrid3 weeks agoEngineering - LI
senior AWS DevOps Architect
NewLibsys, Inc.
Houston, TX🇺🇸On-siteYesterdayShellAWSArgoCD+8Technology