Haystack
← Back to Jobs
Technology

Senior Linux / SRE Infrastructure Engineer

Shree Narayani Networking Solutions LLCChandler, AZ🇺🇸United StatesPosted 10 Aug 2026

Quick Overview

Work Type
Hybrid
Level
Mid Senior

Job Description

Job Summary:

Platform Engineering & Operations
- Administer, monitor, and tune Oracle Enterprise Linux (OEL) environments in large-scale enterprises.
- Oversee design, build, and lifecycle management of Linux servers, storage, and virtualization infrastructure.
- Manage high availability (HA), clustering, and load-balancing for minimal downtime.
- Lead capacity planning and performance optimization initiatives.

Reliability & Automation (SRE Practices)
- Define and implement Site Reliability Engineering (SRE) principles (SLIs, SLOs, error budgets).
- Lead infrastructure automation using tools like Ansible for provisioning, configuration, and patching.
- Build self-healing systems to reduce manual intervention and improve resilience.
- Automate system installation, configuration, and deployment pipelines.

System Administration & Infrastructure Management
- Install, configure, and maintain OEL operating systems and related software.
- Manage Logical Volume Manager (LVM) configurations and distributed file systems.
- Administer network services (DNS, NTP, LDAP/Kerberos, SMTP, OpenSSH) and troubleshoot protocols (TCP/IP, HTTP/S, RPC).

Monitoring, Incident Management & Support
- Implement and enhance system monitoring, alerting, and observability.
- Lead incident response, root cause analysis, and postmortem reviews.
- Drive continuous improvement and oversee break/fix operations.

Security & Compliance
- Ensure systems are secure, hardened, and compliant with security standards.
- Manage patching, vulnerability remediation, and OS upgrades.
- Collaborate with security teams on access control, auditing, and encryption best practices.

Leadership & Collaboration
- Provide technical leadership and mentorship to SRE and infrastructure teams.
- Collaborate with application, DevOps, and platform teams for system reliability.
- Define and enforce operational standards, runbooks, and best practices.

Documentation & Governance
- Maintain comprehensive documentation for architecture and operational procedures.
- Ensure compliance with change management and incident governance frameworks.
- Standardize operational workflows across environments.

Required Skills & Experience
- 5+ years’ Linux system administration in enterprise settings.
- Strong expertise in Oracle Enterprise Linux (OEL) and FPP.
- Experience in high availability systems, virtualization, and storage management.
- Hands-on with automation/configuration tools (Ansible preferred).
- Proficient in scripting/programming (Bash, Python preferred).
- Strong troubleshooting, performance tuning, and incident management skills.
- Solid understanding of enterprise compute, storage, and networking.
- Excellent analytical, problem-solving, communication, and collaboration skills.

Skills

Oracle
Encryption
TCP/IP
Ansible
Bash
DNS
HTTP
LDAP
Python

Similar jobs