Haystack
← Back to Jobs
Full time
Technology

Lead SRE

CoffeeBeansAssamIndiaPosted 29 Jul 2026

Quick Overview

Work Type
Hybrid
Schedule
Full Time
Level
Mid Senior

Job Description

Experience: 7 - 9 Years

Location: Bangalore/Hyderabad/Chennai/Gurgoan/Coimbatore

WorkMode: WFO


Lead SRE / Infrastructure DevOps Engineer (Azure & Oracle Platform)

Key Responsibilities


Design, deploy, and manage cloud infrastructure on Microsoft Azure.

Build and maintain Infrastructure as Code (IaC) using Terraform, ARM Templates, or Bicep to provision and manage infrastructure consistently.

Develop and maintain CI/CD pipelines using Azure DevOps and GitHub Actions to automate infrastructure and application deployments.

Administer Azure core services including Virtual Machines, Virtual Networks, Load Balancers, Azure Kubernetes Service (AKS), App Services, Function Apps, Storage Accounts, Azure SQL, and Azure Monitor.

Implement Infrastructure-as-Code, configuration management, and platform automation using PowerShell, Python, Bash, or Go.

Design, implement, and maintain secure cloud environments using Azure RBAC, Managed Identities, Key Vault, Private Endpoints, Azure Policies, and governance frameworks.

Monitor platform health, infrastructure performance, availability, and security using Azure Monitor, Log Analytics, Application Insights, Microsoft Sentinel, and enterprise monitoring tools.

Manage Oracle database platforms from an infrastructure and platform engineering perspective, including installation, patching coordination, backup validation, high availability, disaster recovery, storage optimization, and capacity planning.

Collaborate with database administrators to identify and resolve Oracle database performance bottlenecks related to infrastructure, operating systems, storage, networking, and resource utilization.

Support Oracle performance tuning by analyzing AWR reports, wait events, execution plans, memory utilization, I/O bottlenecks, and database infrastructure metrics.

Automate routine infrastructure and database platform operations to improve operational efficiency and reliability.

Troubleshoot production incidents across infrastructure, cloud, middleware, and Oracle database platforms while driving root cause analysis and preventive improvements.

Work closely with Architecture, Security, Application, and Operations teams to deliver secure, scalable, and resilient platform solutions.

Create and maintain technical documentation, operational runbooks, architecture diagrams, and standard operating procedures.




Required Skills

Infrastructure & Cloud


Strong hands-on experience with Microsoft Azure infrastructure and cloud-native services.

Experience managing enterprise infrastructure including networking, compute, storage, identity, and platform services.

Strong understanding of Azure governance, subscriptions, resource groups, policies, and management groups.


DevOps & Automation


Strong experience with Terraform, ARM Templates, or Bicep.

Experience implementing CI/CD pipelines using Azure DevOps or GitHub Actions.

Strong scripting skills using PowerShell, Python, Bash, or Go.

Experience with Git, automation frameworks, and Infrastructure as Code.

Good understanding of container platforms such as Docker and Kubernetes (AKS preferred).


Oracle Platform


Good understanding of Oracle Database architecture and administration concepts.

Experience supporting Oracle database infrastructure in production environments.

Ability to identify Oracle database performance bottlenecks involving CPU, memory, storage, network, or SQL execution.

Experience working with Oracle performance tools such as AWR, ADDM, ASH, OEM, and execution plans.

Familiarity with Oracle backup and recovery, RMAN, Data Guard, and high availability concepts.

Ability to work effectively with Oracle DBAs during performance tuning and production troubleshooting.


Security


Experience implementing Azure RBAC, Managed Identities, Key Vault, Private Endpoints, Defender for Cloud, and Azure Policy.

Good understanding of cloud security, governance, compliance, and enterprise security best practices.


Monitoring & Reliability


Experience with Azure Monitor, Log Analytics, Application Insights, Microsoft Sentinel, Prometheus, Grafana, or similar enterprise monitoring platforms.

Strong troubleshooting skills across infrastructure, cloud, operating systems, middleware, and database platforms.

Understanding of SRE principles including observability, incident management, automation, reliability engineering, and operational excellence.



Qualifications


Bachelor's degree in Computer Science, Information Technology, Engineering, or a related discipline.

7+ years of experience in Infrastructure Engineering, DevOps, Platform Engineering, Cloud Engineering, or Site Reliability Engineering.

Minimum 3+ years of hands-on experience with Microsoft Azure.

Experience supporting enterprise Oracle database platforms in production environments.

Holds any two Microsoft Azure Certifications such as:

AZ-104: Microsoft Azure Administrator Associate

AZ-305: Azure Solutions Architect Expert

AZ-400: Azure DevOps Engineer Expert

AZ-500: Azure Security Engineer Associate

AZ-700: Azure Network Engineer Associate



Key Competencies


Strong Infrastructure Engineering mindset

DevOps and Automation-first approach

Platform Reliability and Operational Excellence

Oracle Infrastructure Performance Troubleshooting

Production Incident Management and Root Cause Analysis

Cloud Security and Governance

Excellent analytical and problem-solving skills

Strong collaboration and stakeholder management

Ownership, accountability, and continuous improvement

Passion for automation, platform engineering, and cloud innovation

Skills

Docker
Oracle
SQL
Azure
Bash
Git
GitHub Actions
Grafana
Kubernetes
PowerShell
Prometheus
Python
Stakeholder Management
Terraform
Vault

Similar jobs