Haystack
← Back to Jobs
Technology

Lead Principal Site Reliability Engineer (SRE)

Oracle CorporationUnited States🇺🇸United StatesPosted 15 Aug 2026

Quick Overview

Salary
$104.6k - $264.1k/yr
Work Type
Hybrid
Level
Leader

Job Description

Job Description

Location / Work Authorization / Clearance
  • Role is based in the United States.
  • U.S. citizenship required due to security clearance requirements.
  • No visa sponsorship available.
  • Must be able to obtain and maintain the required security clearance.

We are seeking a highly motivated Lead Principal Site Reliability Engineer (SRE) to support production cloud platforms and mission-critical applications across commercial and sovereign cloud environments. This role combines deep systems engineering expertise with software engineering principles to build resilient, secure, and highly automated infrastructure at scale. As a Senior SRE, you will own the operational health of critical cloud services by improving reliability, automating operational processes, reducing toil, and driving continuous improvements across the service lifecycle. You will partner closely with software engineering, cloud infrastructure, networking, security, database, and operations teams to ensure services meet demanding availability, performance, scalability, and security objectives.
This role requires a strong operational mindset, technical leadership, and the ability to troubleshoot complex distributed systems while designing long-term engineering solutions that eliminate recurring operational issues.

Responsibilities

Reliability Engineering & Operations
  • Own the reliability, availability, scalability, and performance of production cloud services.
  • Operate and support production environments, including compute, storage, networking, databases, and cloud-native services.
  • Maintain service health through proactive monitoring, observability, capacity planning, and performance optimization.
  • Define, measure, and improve Service Level Indicators (SLIs), Service Level Objectives (SLOs), and error budgets.
  • Ensure disaster recovery, backup, restore, and business continuity capabilities are implemented, tested, and maintained.
  • Lead production readiness reviews and ensure operational excellence across all supported services.

Automation & Engineering Excellence
  • Design, develop, and maintain automation that eliminates manual operational work and improves reliability.
  • Build Infrastructure as Code (IaC) solutions using Terraform and configuration management tools.
  • Develop automation using Python, Go, Bash, or similar languages.
  • Build and enhance CI/CD pipelines that support safe, repeatable deployments.
  • Continuously identify and eliminate operational toil and technical debt.
  • Improve operational efficiency through intelligent automation and orchestration.

Incident Response & Problem Management
  • Lead response efforts during high-severity production incidents.
  • Serve as the technical lead during incident bridges and major outages.
  • Perform deep technical investigations to identify root causes of complex production issues.
  • Conduct Root Cause Analysis (RCA) and implement corrective and preventive actions.
  • Develop long-term engineering solutions that prevent recurrence of operational issues.
  • Coordinate complex cross-functional incident response involving engineering, infrastructure, networking, security, and external vendors.

Cloud Infrastructure & Platform Engineering
  • Support large-scale cloud infrastructure across Oracle Cloud Infrastructure (OCI) and cloud-native platforms.
  • Manage Linux-based production environments.
  • Support Kubernetes clusters, containerized workloads, and cloud-native applications.
  • Optimize compute, storage, networking, and database performance.
  • Perform upgrades, patching, maintenance, and lifecycle management while maintaining high availability.
  • Lead capacity forecasting and infrastructure planning to meet future business demand.

Observability & Performance Engineering
  • Design and improve observability using metrics, logs, traces, dashboards, and alerting.
  • Build operational dashboards that provide actionable service health insights.
  • Analyze service behavior under load and optimize system performance.
  • Develop proactive monitoring strategies to detect issues before customer impact.
  • Leverage operational analytics to drive continuous service improvements.

Identity & Security Engineering
  • Provide deep technical expertise in identity architecture and authorization.
  • Design and support Zero Trust architectures, service identity, authentication, federation, and authorization models.
  • Ensure services comply with corporate, regulatory, and federal security standards.
  • Drive vulnerability remediation and security hardening across production environments.
  • Partner closely with cybersecurity teams to improve cloud security posture.

AI-Assisted Engineering & Intelligent Automation
  • Utilize AI-assisted engineering tools such as GitHub Copilot, Codex, Cursor, Claude Code, or similar technologies to improve engineering productivity.
  • Apply Large Language Models (LLMs) to troubleshooting, operational automation, incident management, knowledge management, and workflow optimization.
  • Explore agentic workflows and intelligent automation frameworks to improve operational efficiency and service reliability.
  • Evaluate emerging AI technologies that improve engineering effectiveness and operational excellence.

Technical Leadership
  • Act as the primary technical escalation point for complex production issues.
  • Mentor junior engineers and provide technical leadership across the organization.
  • Lead design discussions and operational reviews.
  • Drive adoption of SRE best practices throughout engineering organizations.
  • Develop operational documentation, runbooks, and standard operating procedures.
  • Communicate effectively with engineering teams, leadership, customers, and executive stakeholders.

Required Qualifications
  • Bachelor's degree in Computer Science, Engineering, Information Technology, or a related technical discipline (or equivalent practical experience).
  • 8+ years of experience in Site Reliability Engineering, DevOps, Cloud Engineering, Systems Engineering, or Production Operations.
  • Experience supporting large-scale production environments with strict availability and uptime requirements.
  • Strong Linux/Unix systems administration experience.
  • Strong programming and scripting skills in Python, Go, Java, Bash, or similar languages.
  • Experience developing Infrastructure as Code using Terraform or equivalent technologies.
  • Experience building CI/CD pipelines using Jenkins, GitHub Actions, GitLab CI, Azure DevOps, or similar platforms.
  • Hands-on experience with Oracle Cloud Infrastructure (OCI), AWS, Azure, or Google Cloud Platform.
  • Experience with Kubernetes, Docker, and container orchestration platforms.
  • Strong understanding of distributed systems architecture and cloud-native applications.
  • Experience with observability platforms such as Prometheus, Grafana, ELK/OpenSearch, Splunk, Datadog, New Relic, or OCI Observability services.
  • Strong understanding of networking fundamentals, including TCP/IP, DNS, load balancing, routing, APIs, service endpoints, and network security.
  • Experience participating in and leading major incident response and operational bridges.
  • Strong troubleshooting and root cause analysis skills across complex distributed systems.
  • Experience with disaster recovery, backup strategies, and business continuity planning.
  • Excellent written and verbal communication skills.

Qualifications

Disclaimer:

Certain U.S. based or U.S. customer or client-facing roles may be required to comply with applicable requirements, such as immunization/occupational health mandates, and/or drug testing requirements.

Range and benefit information provided in this posting are specific to the stated locations only

US: Hiring Range in USD from: $104,600 to $264,100 per annum. May be eligible for bonus, equity, and compensation deferral.

Oracle maintains broad salary ranges for its roles in order to account for variations in knowledge, skills, experience, market conditions and locations, as well as reflect Oracle's differing products, industries and lines of business.
Candidates are typically placed into the range based on the preceding factors as well as internal peer equity.

Oracle US offers a comprehensive benefits package which includes the following:
1. Medical, dental, and vision insurance, including expert medical opinion
2. Short term disability and long term disability
3. Life insurance and AD&D
4. Supplemental life insurance (Employee/Spouse/Child)
5. Health care and dependent care Flexible Spending Accounts
6. Pre-tax commuter and parking benefits
7. 401(k) Savings and Investment Plan with company match
8. Paid time off: Flexible Vacation is provided to all eligible employees assigned to a salaried (non-overtime eligible) position. Accrued Vacation is provided to all other employees eligible for vacation benefits. For employees working at least 35 hours per week, the vacation accrual rate is 13 days annually for the first three years of employment and 18 days annually for subsequent years of employment. Vacation accrual is prorated for employees working between 20 and 34 hours per week. Employees working fewer than 20 hours per week are not eligible for vacation.
9. 11 paid holidays
10. Paid sick leave: 72 hours of paid sick leave upon date of hire. Refreshes each calendar year. Unused balance will carry over each year up to a maximum cap of 112 hours.
11. Paid parental leave
12. Adoption assistance
13. Employee Stock Purchase Plan
14. Financial planning and group legal
15. Voluntary benefits including auto, homeowner and pet insurance

The role will generally accept applications for at least three calendar days from the posting date or as long as the job remains posted.
Career Level - IC5

About Us

Only Oracle brings together the data, infrastructure, applications, and expertise to power everything from industry innovations to life-saving care. And with AI embedded across our products and services, we help customers turn that promise into a better future for all. Discover your potential at a company leading the way in AI and cloud solutions that impact billions of lives.

True innovation starts when everyone is empowered to contribute. That's why we're committed to growing a workforce that promotes opportunities for all with competitive benefits that support our people with flexible medical, life insurance, and retirement options. We also encourage employees to give back to their communities through our volunteer programs.

We're committed to including people with disabilities at all stages of the employment process. If you require accessibility assistance or accommodation for a disability at any point, let us know by emailing or by calling 1- in the United States.

Oracle is an Equal Employment Opportunity Employer. All qualified applicants will receive consideration for employment without regard to race, color, religion, sex, national origin, sexual orientation, gender identity, disability and protected veterans' status, or any other characteristic protected by law. Oracle will consider for employment qualified applicants with arrest and conviction records pursuant to applicable law.

Skills

Docker
Oracle
AWS
ELK
Load Balancing
New Relic
Splunk
TCP/IP
Azure
Bash
DNS
Datadog
GitHub Actions
GitLab CI
Google Cloud
Grafana
Java
Jenkins
Kubernetes
Prometheus
Python
SAFe
Terraform
Zero Trust

Similar jobs