Senior DevOps Engineer – HPC / EDA / SLURM
Quick Overview
Job Description
Senior DevOps Engineer – HPC / EDA / SLURM
Location: Remote
Duration: 3 Months
Experience: 5+ Years
Employment Type: Contract
About the Role
We are seeking a Senior DevOps Engineer to support a High Performance Computing (HPC) and Electronic Design Automation (EDA) infrastructure team on a 3-month contract engagement. This role will work closely with the IT Datacenter organization and requires a highly experienced engineer who can operate independently with minimal ramp-up time.
The ideal candidate will have strong hands-on expertise in Linux HPC environments, SLURM, Ansible, infrastructure automation, enterprise identity management, storage, and datacenter migrations.
Key Responsibilities
HPC / EDA Platform Operations
- Administer and support SLURM-based HPC compute environments, including partition configuration and migration planning.
- Plan and execute HPC/EDA compute and storage infrastructure migrations across datacenters.
- Develop migration strategies, assess risks and dependencies, and define operational trade-offs.
- Create formal Method of Procedure (MOP) documents, implementation plans, and operational runbooks.
- Coordinate with EDA/SPE, storage, identity, and infrastructure teams to execute platform changes.
- Validate storage volumes, application access, authentication, and service continuity following migrations.
- Define HPC storage service tiers and gather performance and capacity requirements for EDA workloads.
Linux Systems Engineering
- Administer production SUSE Linux Enterprise Server (SLES) 12 and SLES 15 environments.
- Build and maintain custom Linux OS images and installation media using Kiwi NG.
- Support bare-metal provisioning using RackN / Digital Rebar Provision and PXE-based deployment.
- Provision and configure VMware vSphere virtual machines supporting HPC services.
- Troubleshoot Linux services, system daemons, operational scripts, and production issues.
Automation & Infrastructure as Code
- Develop and maintain Ansible playbooks and roles for Linux configuration, authentication, and platform automation.
- Ensure compatibility across multiple SLES versions.
- Manage infrastructure code through Git/GitHub and maintain artifacts through Artifactory.
- Participate in pull requests, code reviews, and inner-source infrastructure development.
- Manage production changes through ServiceNow change-management processes.
Identity & Access Management
- Configure and integrate enterprise identity technologies including:
- Okta
- Active Directory
- LDAP
- NIS
- VAS
- SSSD
- Audit and reconcile Linux user/group information, including UID/GID mappings.
- Troubleshoot authentication and access issues across HPC compute and storage environments.
- Extend SSSD-based corporate authentication to new compute environments.
Monitoring & Operational Support
- Evaluate and implement log-management solutions, including Splunk integration.
- Troubleshoot production services such as VNC, NIS, AutoFS, Zabbix, and related Linux services.
- Develop technical documentation, architecture diagrams, implementation guides, and end-user documentation using Confluence.
Required Qualifications
- 5+ years of experience in DevOps, Platform Engineering, Linux Systems Engineering, or a related role.
- Hands-on experience administering HPC clusters using SLURM or comparable workload managers.
- Experience supporting EDA, scientific computing, semiconductor, or high-performance engineering environments.
- Strong production experience with Ansible, including playbook and role development.
- Experience with bare-metal provisioning tools such as RackN / Digital Rebar, Cobbler, or equivalent.
- Proven experience planning and executing datacenter or infrastructure migrations.
- Strong understanding of enterprise Linux authentication and identity technologies, including SSSD, LDAP, Active Directory, NIS, or Okta.
- Experience with NetApp or comparable enterprise storage platforms in HPC environments.
- Ability to create detailed MOPs, runbooks, architecture diagrams, and technical documentation.
- Strong communication skills and ability to coordinate across multiple technical teams.
Required Technical Skills
| Category | Required Skills |
|---|---|
| HPC / EDA | SLURM, HPC compute & storage, EDA infrastructure, datacenter migration |
| Linux | SLES 12/15, Linux services, ESXi 8.0, Kiwi NG, dracut |
| Automation | Ansible, YAML, Python, Bash, Perl |
| Provisioning | RackN / Digital Rebar Provision, PXE |
| Virtualization | VMware vSphere |
| Identity | SSSD, Okta, AD, LDAP, NIS, VAS |
| Storage | NetApp SVM, NFS, AutoFS, RootSquash, IOPS/capacity planning |
| DevOps | Git, GitHub, Artifactory |
| Monitoring | Splunk, Zabbix, Linux log management |
| ITSM / Documentation | ServiceNow, Jira, Confluence, MOPs, technical diagrams |
Preferred Qualifications
- Enterprise experience with SLES 12 and SLES 15.
- Hands-on experience with RackN / Digital Rebar Provision.
- Experience creating custom OS images using Kiwi NG or similar tools.
- VMware vSphere experience supporting HPC infrastructure.
- Experience migrating large configuration artifacts and binaries to Artifactory.
- Background in semiconductor, storage, EDA, or high-tech manufacturing IT environments.
- Experience with enterprise-scale infrastructure modernization and migration programs.
Skills
Similar jobs
DevOps Engineer
Innova Solutions, Inc · Greenwood Village, United States
12 minutes ago$60 - $70/hrData Infrastructure Site Reliability Engineer
Marici Solutions · United States
16 minutes agoSite Reliability / Environment Support Lead
isolve technology inc · United States
16 minutes agoSITEC - Cloud Platform Engineer - MacDill AFB with Security Clearance
Peraton · Macdill AFB, United States
38 minutes ago$104k - $166k/yrDevOps Software Engineer - TS/SCI Cleared
Leidos · Rockville, United States
39 minutes ago$107.9k - $195.1k/yrDevOps Software Engineer - TS/SCI Cleared
Leidos · Silver Spring, United States
39 minutes ago$107.9k - $195.1k/yr