Haystack
← Back to Jobs
Other
IS

Heavy AWS + HPC + Parallel Cluster + Slrum

Infodyne SolutionsUnited States🇺🇸United StatesPosted Sep 22, 2026

Quick Overview

Seniority
Mid Senior
Work mode
Hybrid
Location
United States
Posted
Yesterday
DockerAWSMachine LearningAgileAzureBashCloudFormationGoogle CloudHadoopKubernetesPyTorchPythonTensorFlowTerraform

Job Description

Heavy AWS + HPC High Performance Computing, Parallel Cluster * Slrum is a must

Location: US, Remote

We're seeking a DevOps Engineer with High Performance Computing (HPC) expertise to design, automate, and maintain the infrastructure supporting the compute-intensive scientific and research workloads (e.g., genomics, molecular modeling, drug discovery simulations). This role bridges traditional DevOps practices with specialized HPC cluster management in a life sciences/pharma environment.

Cloud Systems Engineer (AI/ML & HPC Specialization)

Key Responsibilities:

  • Design, implement, and manage cloud-based infrastructure that supports AI/ML workflows. for
  • Collaborate with data scientists and ML engineers to deploy scalable machine learning models into production.
  • Ensure the security, scalability, and reliability of AI/ML systems in the cloud.
  • Optimize cloud resources for cost-effective and efficient use.
  • Stay current with the latest in cloud services, AI/ML tools, and industry best practices.
  • Provide technical leadership and guidance in cloud and AI/ML architecture.
  • Develop and maintain CI/CD pipelines for AI/ML model training and deployment.
  • Monitor and troubleshoot AI/ML applications and cloud environments.
  • Document system design and operational procedures.
  • Collaborate with AI/ML and HPC teams to understand their computing and storage needs.

Qualifications:

  • Bachelor s or Master s degree in Computer Science, Engineering, or related field.
  • Proven experience in cloud computing (AWS, Azure, Google Cloud Platform) and cloud architecture.
  • Strong background in AI/ML technologies, with experience in deploying ML models.
  • Proficiency in scripting languages (Python, Bash) and containerization technologies (Docker, Kubernetes).
  • Proficiency with virtual compute environments (EC2).
  • Hands-on experience with High Performance Computing (HPC) and server node Cluster Management
  • Strong Knowledge of Linux/Unix operating systems (RHEL/Ubuntu)
  • Experience with job schedulers (like SLURM, PBS), resource management, and system monitoring tools (DynaTrace).
  • Understanding of storage solutions and file systems used in HPC (such as Lustre, GPFS).
  • Experience with infrastructure as code (IaC) tools like Terraform or CloudFormation.
  • Knowledge of networking, security, and database technologies in a cloud environment.
  • Excellent problem-solving, communication, and team collaboration skills.

Preferred Skills:

  • Familiarity with machine learning frameworks (TensorFlow, PyTorch) and data pipelines.
  • Certifications in cloud architecture (AWS Certified Solutions Architect, Google Cloud Professional Cloud Architect, etc.).
  • Experience in an Agile development environment.
  • Prior work with distributed computing and big data technologies (Hadoop, Spark).
  • Operational experience running large scale platforms, including AI/ML platforms

Similar jobs