Quick Overview
Seniority
Mid Senior
Work mode
Hybrid
Location
United States
Posted
Yesterday
DockerAWSMachine LearningAgileAzureBashCloudFormationGoogle CloudHadoopKubernetesPyTorchPythonTensorFlowTerraform
Job Description
Heavy AWS + HPC High Performance Computing, Parallel Cluster * Slrum is a must
Location: US, Remote
We're seeking a DevOps Engineer with High Performance Computing (HPC) expertise to design, automate, and maintain the infrastructure supporting the compute-intensive scientific and research workloads (e.g., genomics, molecular modeling, drug discovery simulations). This role bridges traditional DevOps practices with specialized HPC cluster management in a life sciences/pharma environment.
Cloud Systems Engineer (AI/ML & HPC Specialization)
Key Responsibilities:
- Design, implement, and manage cloud-based infrastructure that supports AI/ML workflows. for
- Collaborate with data scientists and ML engineers to deploy scalable machine learning models into production.
- Ensure the security, scalability, and reliability of AI/ML systems in the cloud.
- Optimize cloud resources for cost-effective and efficient use.
- Stay current with the latest in cloud services, AI/ML tools, and industry best practices.
- Provide technical leadership and guidance in cloud and AI/ML architecture.
- Develop and maintain CI/CD pipelines for AI/ML model training and deployment.
- Monitor and troubleshoot AI/ML applications and cloud environments.
- Document system design and operational procedures.
- Collaborate with AI/ML and HPC teams to understand their computing and storage needs.
Qualifications:
- Bachelor s or Master s degree in Computer Science, Engineering, or related field.
- Proven experience in cloud computing (AWS, Azure, Google Cloud Platform) and cloud architecture.
- Strong background in AI/ML technologies, with experience in deploying ML models.
- Proficiency in scripting languages (Python, Bash) and containerization technologies (Docker, Kubernetes).
- Proficiency with virtual compute environments (EC2).
- Hands-on experience with High Performance Computing (HPC) and server node Cluster Management
- Strong Knowledge of Linux/Unix operating systems (RHEL/Ubuntu)
- Experience with job schedulers (like SLURM, PBS), resource management, and system monitoring tools (DynaTrace).
- Understanding of storage solutions and file systems used in HPC (such as Lustre, GPFS).
- Experience with infrastructure as code (IaC) tools like Terraform or CloudFormation.
- Knowledge of networking, security, and database technologies in a cloud environment.
- Excellent problem-solving, communication, and team collaboration skills.
Preferred Skills:
- Familiarity with machine learning frameworks (TensorFlow, PyTorch) and data pipelines.
- Certifications in cloud architecture (AWS Certified Solutions Architect, Google Cloud Professional Cloud Architect, etc.).
- Experience in an Agile development environment.
- Prior work with distributed computing and big data technologies (Hadoop, Spark).
- Operational experience running large scale platforms, including AI/ML platforms
Similar jobs
- RE
Field Marketer
NewRenuity
Minneapolis, Minnesota🇺🇸$55k - $75k/yrOn-site26 minutes agoCRMLead GenerationOutreach - OF
Consumer Lending Advisor
NewOneMain Financial
Livonia, MI🇺🇸On-site14 hours agoBusiness Development - VS
Sr. Technical Lead
NewVG Systems
Huntsville, Alabama🇺🇸Remote28 minutes agoMicrosoft OfficePortfolio Management - MA
AI Technical Lead
NewMesa Associates, Inc.
Madison, Alabama🇺🇸Hybrid28 minutes agoSQLMLOpsMachine Learning+4 - CO
Technical Deployment Lead - East
NewCoreWeave
Marble, North Carolina🇺🇸$90k - $102k/yrOn-site28 minutes agoSpringComplianceRoot Cause Analysis+1 - NC
Technical Lead (1) - C Shift - 7:30pm - 7:45am - (4 on 4 off)
NewNGK Ceramics USA, Inc.
Mooresville, North Carolina🇺🇸Hybrid28 minutes ago5SComplianceConcrete+2