Haystack
← Back to Jobs
Remote
Technology
SI

Senior HPC DevOps Engineer

Stellent IT LLCRancho Cordova, CA🇺🇸United StatesPosted Oct 7, 2026

Quick Overview

Seniority
Mid Senior
Work mode
Remote
Location
Rancho Cordova, CA, United States
Posted
21 hours ago
SplunkActive DirectoryAnsibleAzureBashConfluenceCordovaDatadogGitLDAPPythonTerraform

Job Description

Senior DevOps Engineer HPC / EDA / SLURM / Azure

Location: Remote Rancho Cordova, CA
Duration: 12 months

Job Description

Role Summary:
Client is seeking a Senior DevOps Engineer to support its HPC and EDA cloud infrastructure. The role requires strong hands-on experience with Linux HPC environments, SLURM, Ansible, Terraform, Azure, identity management, and enterprise storage. The candidate should be able to work independently and support production infrastructure with minimal ramp-up.

Key Responsibilities

  • Administer SLURM-based HPC clusters, compute/storage environments, and EDA infrastructure.
  • Develop and maintain Ansible playbooks/roles and Terraform automation.
  • Support Linux environments including SLES 15 and Ubuntu.
  • Manage enterprise authentication using SSSD, LDAP, Active Directory, and Okta.
  • Support NetApp/NFS, AutoFS, RootSquash, storage capacity and IOPS planning.
  • Manage Azure EDA user environments, including ThinLinc/VNC.
  • Use Git/GitHub, Artifactory and participate in code reviews/PRs.
  • Implement monitoring/logging using Splunk and Datadog.
  • Troubleshoot production Linux services and operational issues.
  • Manage infrastructure changes through ServiceNow.
  • Create MOPs, runbooks, architecture diagrams, implementation guides, and Confluence documentation.
  • Coordinate infrastructure changes with EDA, NAND, storage, and IAM teams.

Must-Haves

  • 5+ years in DevOps, Platform Engineering, or Linux Systems Engineering.
  • Strong hands-on HPC cluster administration experience.
  • SLURM or equivalent workload manager experience.
  • Ansible + Terraform production automation.
  • Strong Linux/SLES/Ubuntu experience.
  • Experience with Azure cloud compute.
  • SSSD, LDAP, AD, Okta / enterprise Linux authentication.
  • NetApp/NFS or comparable enterprise HPC storage.
  • Experience supporting EDA/scientific computing environments.
  • Strong Python/Bash/YAML scripting skills.
  • Experience with Git/GitHub and ServiceNow.
  • Strong technical documentation and cross-team communication skills.

Nice-to-Haves

  • SLES 12/15 enterprise experience.
  • Artifactory migration experience.
  • Semiconductor, storage, or high-tech manufacturing IT background.

Must-Haves

1. 5+ years of DevOps / Platform Engineering / Linux Systems Engineering experience.

2. Strong HPC cluster administration experience.

3. Hands-on SLURM workload manager experience.

4. Experience supporting EDA or scientific computing environments.

5. Strong Ansible experience - playbooks, roles, and production automation.

6. Strong Terraform / Infrastructure as Code experience.

7. Advanced Linux skills, preferably SLES 15 and/or Ubuntu.

8. Hands-on Azure cloud compute experience.

9. Enterprise Linux identity/authentication experience with SSSD, LDAP, Active Directory, and/or Okta.

10. NetApp / NFS or comparable enterprise storage experience in HPC environments.

11. Strong scripting skills with Python and Bash; YAML/Ansible.

12. Git/GitHub experience, including PRs/code reviews.

13. Experience with ServiceNow change management.

14. Ability to create MOPs, runbooks, architecture diagrams, and technical documentation.

15. Strong communication and ability to work independently with multiple infrastructure/engineering teams.

Navnish kumar

Sr. IT Technical Recruiter

Stellent IT Phone :

Email: navnish
Gtalk: navnish om

Similar jobs