Quick Overview
Job Description
๐ง๐ต๐ถ๐ ๐ฟ๐ผ๐น๐ฒ ๐ถ๐ ๐ณ๐ผ๐ฟ ๐ผ๐ป๐ฒ ๐ผ๐ณ ๐๐ต๐ฒ ๐ช๐ฒ๐ฒ๐ธ๐ฑ๐ฎ๐'๐ ๐ฐ๐น๐ถ๐ฒ๐ป๐๐
๐ฆ๐ฎ๐น๐ฎ๐ฟ๐ ๐ฟ๐ฎ๐ป๐ด๐ฒ: ๐ฅ๐ ๐ฎ๐ฌ๐ฌ๐ฌ๐ฌ๐ฌ๐ฌ - ๐ฅ๐ ๐ฏ๐ฌ๐ฌ๐ฌ๐ฌ๐ฌ๐ฌ (๐ถ๐ฒ ๐๐ก๐ฅ ๐ฎ๐ฌ-๐ฏ๐ฌ ๐๐ฃ๐)
Experience: 5+ yrs
Location: Bengaluru, Karnataka, India, Hyderabad, Telangana, India
Job Type: Full-time
We are looking for an experiencedย Linux SME / SRE Engineerย with strong expertise inย Core Linux Administration, RHEL, and PCS/Pacemaker clusteringย to support business-critical production environments.
The role focuses on maintaining highly available Linux infrastructure, resolving complex production issues, ensuring system reliability, and supporting clustered environments. The ideal candidate will have strong hands-on troubleshooting capabilities, a solid understanding of high-availability architectures, and the ability to work effectively with clients and technical stakeholders.
Key Responsibilities
- Administer and supportย Linux-based production environmentsย across critical infrastructure.
- Perform day-to-dayย Core Linux administration, configuration, monitoring, maintenance, and troubleshooting.
- Manage, monitor, configure, and troubleshootย PCS/Pacemaker high-availability clusters.
- Ensure availability, reliability, stability, and performance of Linux infrastructure and clustered services.
- Troubleshoot complex and critical production incidents and drive issues through to resolution.
- Perform root-cause analysis and implement sustainable solutions for recurring infrastructure problems.
- Monitor system and cluster health and proactively identify potential availability or performance issues.
- Support failover, recovery, maintenance, and operational activities across high-availability environments.
- Collaborate with clients, infrastructure teams, application teams, and other technical stakeholders on incidents and enhancements.
- Participate in incident management, problem management, change management, and production maintenance activities.
- Follow SRE practices for monitoring, reliability improvement, incident response, and operational efficiency.
- Maintain technical documentation, operational procedures, troubleshooting guides, and support records.
- Participate in rotational shifts to provide continuous production support.
- Identify opportunities to automate repetitive infrastructure tasks and improve operational efficiency.
- Support infrastructure changes, upgrades, patching, and maintenance activities in accordance with established processes.
- Contribute to service reliability, availability, and continuous improvement initiatives.
What Makes You a Great Fit
- 5โ9 years of overall experienceย in Linux administration, infrastructure engineering, SRE, or production support, with a maximum of 10 years preferred.
- Minimumย 4 years of hands-on experience with PCS/Pacemaker cluster administration.
- Strong expertise inย Core Linux Administrationย and production infrastructure support.
- Strong hands-on experience withย RHEL (Red Hat Enterprise Linux).
- Solid understanding ofย High Availability, clustering, failover, resource management, and cluster troubleshooting.
- Proven experience supportingย critical production environmentsย with strict availability and reliability requirements.
- Strong troubleshooting, debugging, root-cause analysis, and incident-resolution capabilities.
- Experience working with production monitoring, incident management, and infrastructure maintenance processes.
- Strong understanding ofย SRE and ITIL practicesย is desirable.
- Excellent communication and client-facing skills with the ability to explain technical issues clearly to stakeholders.
- Strong stakeholder-management and collaboration skills.
- Ability to work effectively under pressure during critical production incidents.
- Willingness to work inย rotational shifts, including scheduled production-support coverage.
- Experience withย VMware administrationย is an advantage.
- Exposure toย AWS or other cloud platformsย is desirable.
- Knowledge ofย Oracle Databaseย and its infrastructure dependencies is an advantage.
- Strong ownership mindset with a focus on system reliability, operational excellence, and continuous improvement.
Similar jobs
- WA
Linux SME / SRE Engineer
NewWeekday AI
Bengaluru, Karnataka๐ฎ๐ณOn-site15 hours agoOracleAWSContinuous Improvement+1Technology - MA
Senior DevOps Engineer
NewMarketnode
Bengaluru, Karnataka๐ฎ๐ณHybrid14 hours agoDockerELKLoad Balancing+11Technology - AE
Technical Operations Engineer
NewArbor Education
Thiruvananthapuram, Keralam๐ฎ๐ณOn-site18 hours agoPHPSQLAWS+11Technology - VE
Senior SRE/DevOps Engineer
NewVerve
Bengaluru๐ฎ๐ณHybrid19 hours agoGCPRustCapacity Planning+7Technology - OK
Senior Site Reliability Engineer
NewOkta
Bengaluru๐ฎ๐ณHybrid21 hours agoGCPMySQLSQL+23Technology - MI
Site Reliability Engineer III
NewMitratech
Mitratech India๐ฎ๐ณHybrid17 hours agoDynamoDBPackerAWS+12Technology