Quick Overview
Job Description
𝗧𝗵𝗶𝘀 𝗿𝗼𝗹𝗲 𝗶𝘀 𝗳𝗼𝗿 𝗼𝗻𝗲 𝗼𝗳 𝘁𝗵𝗲 𝗪𝗲𝗲𝗸𝗱𝗮𝘆'𝘀 𝗰𝗹𝗶𝗲𝗻𝘁𝘀
𝗦𝗮𝗹𝗮𝗿𝘆 𝗿𝗮𝗻𝗴𝗲: 𝗥𝘀 𝟮𝟬𝟬𝟬𝟬𝟬𝟬 - 𝗥𝘀 𝟯𝟬𝟬𝟬𝟬𝟬𝟬 (𝗶𝗲 𝗜𝗡𝗥 𝟮𝟬-𝟯𝟬 𝗟𝗣𝗔)
Experience: 5+ yrs
Location: Bengaluru, Karnataka, India, Hyderabad, Telangana, India
Job Type: Full-time
We are looking for an experienced Linux SME / SRE Engineer with strong expertise in Core Linux Administration, RHEL, and PCS/Pacemaker clustering to support business-critical production environments.
The role focuses on maintaining highly available Linux infrastructure, resolving complex production issues, ensuring system reliability, and supporting clustered environments. The ideal candidate will have strong hands-on troubleshooting capabilities, a solid understanding of high-availability architectures, and the ability to work effectively with clients and technical stakeholders.
Key Responsibilities
- Administer and support Linux-based production environments across critical infrastructure.
- Perform day-to-day Core Linux administration, configuration, monitoring, maintenance, and troubleshooting.
- Manage, monitor, configure, and troubleshoot PCS/Pacemaker high-availability clusters.
- Ensure availability, reliability, stability, and performance of Linux infrastructure and clustered services.
- Troubleshoot complex and critical production incidents and drive issues through to resolution.
- Perform root-cause analysis and implement sustainable solutions for recurring infrastructure problems.
- Monitor system and cluster health and proactively identify potential availability or performance issues.
- Support failover, recovery, maintenance, and operational activities across high-availability environments.
- Collaborate with clients, infrastructure teams, application teams, and other technical stakeholders on incidents and enhancements.
- Participate in incident management, problem management, change management, and production maintenance activities.
- Follow SRE practices for monitoring, reliability improvement, incident response, and operational efficiency.
- Maintain technical documentation, operational procedures, troubleshooting guides, and support records.
- Participate in rotational shifts to provide continuous production support.
- Identify opportunities to automate repetitive infrastructure tasks and improve operational efficiency.
- Support infrastructure changes, upgrades, patching, and maintenance activities in accordance with established processes.
- Contribute to service reliability, availability, and continuous improvement initiatives.
What Makes You a Great Fit
- 5–9 years of overall experience in Linux administration, infrastructure engineering, SRE, or production support, with a maximum of 10 years preferred.
- Minimum 4 years of hands-on experience with PCS/Pacemaker cluster administration.
- Strong expertise in Core Linux Administration and production infrastructure support.
- Strong hands-on experience with RHEL (Red Hat Enterprise Linux).
- Solid understanding of High Availability, clustering, failover, resource management, and cluster troubleshooting.
- Proven experience supporting critical production environments with strict availability and reliability requirements.
- Strong troubleshooting, debugging, root-cause analysis, and incident-resolution capabilities.
- Experience working with production monitoring, incident management, and infrastructure maintenance processes.
- Strong understanding of SRE and ITIL practices is desirable.
- Excellent communication and client-facing skills with the ability to explain technical issues clearly to stakeholders.
- Strong stakeholder-management and collaboration skills.
- Ability to work effectively under pressure during critical production incidents.
- Willingness to work in rotational shifts, including scheduled production-support coverage.
- Experience with VMware administration is an advantage.
- Exposure to AWS or other cloud platforms is desirable.
- Knowledge of Oracle Database and its infrastructure dependencies is an advantage.
- Strong ownership mindset with a focus on system reliability, operational excellence, and continuous improvement.