Haystack
← Back to Jobs
Full time
Technology
OR

Principal Site Reliability Engineer

OracleUnited States🇺🇸United StatesPosted Sep 19, 2026

Quick Overview

Seniority
Leader
Employment type
Full Time
Work mode
Hybrid
Location
United States
OracleSQLSQL ServerLoad BalancingActive DirectoryDNSPowerShell

Job Description

hackajob is collaborating with Oracle to connect them with exceptional professionals for this role.
About the team
Your work will help clinicians and healthcare staff access applications consistently, reduce disruption from platform changes, and accelerate recovery when problems occur. You will help the team build a platform that is easier to operate, easier to scale, and more resilient as Oracle Health infrastructure evolves.
Description
About the Opportunity
Help ensure healthcare professionals can reliably access the applications they depend on to deliver patient care.
Oracle Health is seeking a Principal Site Reliability Engineer to strengthen the reliability, performance, security, and scalability of our Citrix platform. The platform delivers access to Oracle Health Millennium and supporting applications used in clinical and business workflows.
In this hands-on role, you will own complex reliability initiatives from design through production operation. You will combine deep Citrix and Windows expertise with software engineering practices to automate operations, improve the user experience, and prevent recurring service failures. You will also guide technical decisions, mentor engineers, and partner across infrastructure, networking, identity, security, database, and application teams.
What You'll Do
  • Own service reliability. Drive improvements across the application-access journey, including authentication, application discovery, session launch, and in-session performance. Define service-level indicators and objectives that reflect customer experience.
  • Engineer resilient Citrix services. Design, operate, and improve Citrix Virtual Apps and Desktops environments, including Delivery Controllers, StoreFront, Virtual Delivery Agents, Provisioning Services, and their supporting dependencies.
  • Automate platform operations. Build maintainable PowerShell scripts and automation for provisioning, configuration, image management, patching, validation, and recovery. Reduce repetitive work through tested, repeatable workflows.
  • Improve observability. Develop actionable monitoring, dashboards, synthetic checks, and alerts for application-launch success, logon performance, session-host registration, provisioning health, and capacity.
  • Lead complex troubleshooting. Coordinate technical response to significant incidents, isolate failures across service dependencies, and drive root-cause analysis and corrective actions that reduce recurrence.
  • Plan for growth and recovery. Assess concurrent-session demand, resource utilization, and capacity headroom. Validate high availability, failover, and disaster recovery through documented exercises.
  • Deliver safe platform changes. Lead Citrix and Windows upgrades, image updates, and cloud migration work, including initiatives involving Oracle Cloud Infrastructure. Establish readiness criteria, deployment validation, and rollback procedures.
  • Strengthen security and configuration practices. Partner with security and infrastructure teams on vulnerability remediation, certificate lifecycle management, access controls, configuration baselines, and drift detection.
  • Provide technical leadership. Own substantial projects independently, explain architectural tradeoffs, review designs and automation, mentor engineers, and improve engineering practices across teams.
  • Participate in the team's production on-call rotation and planned maintenance activities, including work outside standard business hours when required.
Designs and architects infrastructure and service to ensure reliability and functionality. Forecasts demands and responds to capacity needs. Collaborates with software development teams to develop reliable and scalable infrastructures. Exercises judgment when performing data collection to maintain and optimize operations and reliability. Leverages advanced knowledge to perform incident response and/or maintenance tasks. Provides comprehensive health and performance reporting. Identifies and recommends opportunities for automation. Communicates comprehensive information about services and proactively anticipates and articulates the potential impact of changes. Provides comprehensive support for technology and documents incidents. Conducts advanced experiments with new tools and develops and maintains advanced knowledge of site reliability trends.
Responsibilities
Team Required Qualifications
  • Extensive experience in site reliability engineering, systems engineering, infrastructure engineering, or production operations, with demonstrated ownership of critical enterprise services.
  • Deep hands-on experience engineering and supporting production Citrix Virtual Apps and Desktops environments, including session brokering, StoreFront, VDAs, provisioning, and image lifecycle management.
  • Advanced Windows Server troubleshooting skills and a strong understanding of Active Directory, Group Policy, DNS, certificates, and authentication dependencies.
  • Strong PowerShell development skills, including building, testing, troubleshooting, and maintaining operational automation.
  • Ability to diagnose complex issues involving networking, load balancing, compute, storage, operating systems, and application dependencies.
  • Experience with monitoring, incident response, capacity planning, production change management, and high-availability or disaster-recovery practices.
  • A record of leading technical initiatives across teams, making sound decisions with incomplete information, and delivering measurable improvements.
  • Clear written and verbal communication, including technical documentation, incident communication, and mentoring.
Preferred Qualifications
  • Experience supporting Oracle Health Millennium, Cerner applications, or other healthcare application-hosting environments.
  • Experience with Oracle Cloud Infrastructure, cloud migrations, or hybrid infrastructure.
  • Experience with NetScaler ADC/Gateway, Citrix Federated Authentication Service, Citrix Director, profile management, or SQL Server dependencies supporting Citrix.
  • Familiarity with infrastructure as code, configuration management, version control, and automated deployment pipelines.
  • Experience using service-level objectives, error budgets, synthetic monitoring, and resilience testing to guide engineering priorities.
  • Experience improving security and operational reliability in regulated environments.
Org Responsibilities
Capacity Ingestion and Management:
- Designs and architects infrastructure and/or service according to terms for reliability and functionality.
- Forecasts demands for infrastructure and responds to capacity needs, ensuring systems have sufficient resources to handle current and future workloads and identifying resource gaps.
- Collaborates with the software development team to develop infrastructures, ensuring features are reliable and scalable according to deployment requirements.
- Proactively identifies opportunities for prototyping and drives prototyping initiatives (e.g., testing new applications or infrastructures, assisting in onboarding) to explore novel approaches.
Incident and Service Lifecycle Management:
- Exercises judgment when performing data collection, triage, technical analysis, and redirection to maintain and optimize operations and infrastructure reliability.
- Takes proactive steps to monitor services, maintain up-to-date knowledge of their performance, and document their condition.
- Leverages advanced knowledge to perform incident response, root cause analyses, and/or maintenance on assigned services (e.g., software installs, version upgrades, security updates, backup and recovery).
- Provides comprehensive health and performance reporting and takes appropriate actions based on trends in data.
- May perform provisioning to support infrastructure, applications, and services.
- May experiment with new approaches for and performs decommissioning (e.g., shutting down servers, removing data from databases) to remove objects that are no longer needed.
Automation:
- Identifies and recommends opportunities for automation and assesses potential benefits to enhance operational efficiency.
- Develops and implements design, automation tools, or scripts to provide solutions, gather metrics, monitor, analyze, mitigate, or remediate issues/defects within infrastructures.
- Conducts testing on moderately complex automations to ensure they perform tasks correctly and produce expected results.
Technical Communication and Guidance:
- Writes release notes and/or communicates comprehensive information about the scale, capacity, security, performance attributes, and requirements of services and technology with customers and immediate and related teams.
- Proactively anticipates and articulates the potential impact of infrastructure, feature, and tool changes, considering their impact across team operations.
- Serves as a resource to team members on what information to communicate and how to communicate.
Troubleshooting and Resolution:
- Provides comprehensive operational support for technology, serving as a key escalation point for incidents and moderately complex issues arising within Oracle services.
- Drives and actively participates in on-call shifts to address issues.
. click apply for full job details

Similar jobs