← Back to Jobs
Technology
Senior Linux / SRE Infrastructure Engineer
Shree Narayani Networking Solutions LLCChandler, AZ🇺🇸United StatesPosted 10 Aug 2026
Quick Overview
Work Type
Hybrid
Level
Mid Senior
Job Description
Job Summary:
Platform Engineering & Operations
- Administer, monitor, and tune Oracle Enterprise Linux (OEL) environments in large-scale enterprises.
- Oversee design, build, and lifecycle management of Linux servers, storage, and virtualization infrastructure.
- Manage high availability (HA), clustering, and load-balancing for minimal downtime.
- Lead capacity planning and performance optimization initiatives.
Reliability & Automation (SRE Practices)
- Define and implement Site Reliability Engineering (SRE) principles (SLIs, SLOs, error budgets).
- Lead infrastructure automation using tools like Ansible for provisioning, configuration, and patching.
- Build self-healing systems to reduce manual intervention and improve resilience.
- Automate system installation, configuration, and deployment pipelines.
System Administration & Infrastructure Management
- Install, configure, and maintain OEL operating systems and related software.
- Manage Logical Volume Manager (LVM) configurations and distributed file systems.
- Administer network services (DNS, NTP, LDAP/Kerberos, SMTP, OpenSSH) and troubleshoot protocols (TCP/IP, HTTP/S, RPC).
Monitoring, Incident Management & Support
- Implement and enhance system monitoring, alerting, and observability.
- Lead incident response, root cause analysis, and postmortem reviews.
- Drive continuous improvement and oversee break/fix operations.
Security & Compliance
- Ensure systems are secure, hardened, and compliant with security standards.
- Manage patching, vulnerability remediation, and OS upgrades.
- Collaborate with security teams on access control, auditing, and encryption best practices.
Leadership & Collaboration
- Provide technical leadership and mentorship to SRE and infrastructure teams.
- Collaborate with application, DevOps, and platform teams for system reliability.
- Define and enforce operational standards, runbooks, and best practices.
Documentation & Governance
- Maintain comprehensive documentation for architecture and operational procedures.
- Ensure compliance with change management and incident governance frameworks.
- Standardize operational workflows across environments.
Required Skills & Experience
- 5+ years’ Linux system administration in enterprise settings.
- Strong expertise in Oracle Enterprise Linux (OEL) and FPP.
- Experience in high availability systems, virtualization, and storage management.
- Hands-on with automation/configuration tools (Ansible preferred).
- Proficient in scripting/programming (Bash, Python preferred).
- Strong troubleshooting, performance tuning, and incident management skills.
- Solid understanding of enterprise compute, storage, and networking.
- Excellent analytical, problem-solving, communication, and collaboration skills.
Platform Engineering & Operations
- Administer, monitor, and tune Oracle Enterprise Linux (OEL) environments in large-scale enterprises.
- Oversee design, build, and lifecycle management of Linux servers, storage, and virtualization infrastructure.
- Manage high availability (HA), clustering, and load-balancing for minimal downtime.
- Lead capacity planning and performance optimization initiatives.
Reliability & Automation (SRE Practices)
- Define and implement Site Reliability Engineering (SRE) principles (SLIs, SLOs, error budgets).
- Lead infrastructure automation using tools like Ansible for provisioning, configuration, and patching.
- Build self-healing systems to reduce manual intervention and improve resilience.
- Automate system installation, configuration, and deployment pipelines.
System Administration & Infrastructure Management
- Install, configure, and maintain OEL operating systems and related software.
- Manage Logical Volume Manager (LVM) configurations and distributed file systems.
- Administer network services (DNS, NTP, LDAP/Kerberos, SMTP, OpenSSH) and troubleshoot protocols (TCP/IP, HTTP/S, RPC).
Monitoring, Incident Management & Support
- Implement and enhance system monitoring, alerting, and observability.
- Lead incident response, root cause analysis, and postmortem reviews.
- Drive continuous improvement and oversee break/fix operations.
Security & Compliance
- Ensure systems are secure, hardened, and compliant with security standards.
- Manage patching, vulnerability remediation, and OS upgrades.
- Collaborate with security teams on access control, auditing, and encryption best practices.
Leadership & Collaboration
- Provide technical leadership and mentorship to SRE and infrastructure teams.
- Collaborate with application, DevOps, and platform teams for system reliability.
- Define and enforce operational standards, runbooks, and best practices.
Documentation & Governance
- Maintain comprehensive documentation for architecture and operational procedures.
- Ensure compliance with change management and incident governance frameworks.
- Standardize operational workflows across environments.
Required Skills & Experience
- 5+ years’ Linux system administration in enterprise settings.
- Strong expertise in Oracle Enterprise Linux (OEL) and FPP.
- Experience in high availability systems, virtualization, and storage management.
- Hands-on with automation/configuration tools (Ansible preferred).
- Proficient in scripting/programming (Bash, Python preferred).
- Strong troubleshooting, performance tuning, and incident management skills.
- Solid understanding of enterprise compute, storage, and networking.
- Excellent analytical, problem-solving, communication, and collaboration skills.
Skills
Oracle
Encryption
TCP/IP
Ansible
Bash
DNS
HTTP
LDAP
Python
Similar jobs
Site Reliability Engineer
TEKsystems c/o Allegis Group · Danville, United States
23 minutes ago$55 - $65/hrDevOps Cloud Engineer
Judge Group, Inc. · Grandview Heights, United States
23 minutes ago$85 - $90/hrSenior AI Platform Engineer
Motion Recruitment Partners, LLC · Fort Worth, United States
25 minutes agoSenior DevOps Engineer
Infobahn Softworld Inc. · United States
1 hour agoAzure DevOps Architect
VISION INFOTECH INC. · United States
1 hour agoData Platform Engineer
Wise Skulls Corp. · Manor, United States
1 hour ago