Haystack
← Back to Jobs
Technology
CD

AI Infrastructure Engineer

Cloud Destinations LLCUnited States🇺🇸United StatesPosted Sep 28, 2026

Quick Overview

Seniority
Mid Senior
Work mode
Hybrid
Location
United States
Posted
21 hours ago
DockerKubernetesPyTorchTensorFlow

Job Description

Position: AI Infrastructure Engineer

Location: Remote

Hiring Mode: 6+ Months Contract

Job Description:

The AI Infrastructure Operations Engineer will manage, maintain, and optimize artificial intelligence (AI) infrastructure across on-premises and cloud environments. This hands-on role focuses on platform availability, performance, resource utilization, troubleshooting, operational excellence, and team enablement. The engineer will support AI workloads and model deployments while providing practical guidance and training on infrastructure tools and best practices. This position emphasizes reliable operations and knowledge transfer rather than strategic development or workload definition.

Responsibilities:

  • Manage and maintain AI infrastructure to support high availability, reliability, and performance.
  • Implement, administer, and optimize AI operations using NVIDIA Mission Control, Bright Cluster Manager, Run:ai, and related platform tools.
  • Collaborate with cross-functional teams to support AI workloads and improve compute, accelerator, storage, and scheduling resource utilization.
  • Provide training, mentoring, and hands-on guidance to team members on AI infrastructure tools, operational practices, and troubleshooting methods.
  • Monitor platform health and system performance, diagnose issues, and resolve incidents to minimize downtime and improve resource allocation.
  • Assist with the deployment, operation, and scaling of AI models and applications across supported environments.
  • Evaluate relevant advancements in AI infrastructure technology and recommend practical operational improvements.
  • Develop and maintain clear documentation for configurations, procedures, troubleshooting, and infrastructure management best practices.

Required Qualifications:

  • Proven experience managing AI infrastructure and day-to-day platform operations.
  • Hands-on proficiency with NVIDIA Mission Control or Bright Cluster Manager and Run:ai.
  • Proficiency administering Linux operating systems, including Ubuntu and Red Hat Enterprise Linux (RHEL).
  • Strong understanding of high-performance computing (HPC) environments and resource management concepts.
  • Experience supporting both cloud platforms and on-premises infrastructure.
  • Strong troubleshooting, analytical, and problem-solving skills with close attention to detail.
  • Ability to collaborate effectively across technical teams and communicate clearly with varied audiences.
  • Experience training and mentoring technical team members.
  • Bachelor s degree in Computer Science, Engineering, or a related field, or equivalent professional experience.

Preferred Qualifications:

  • Experience with containerization and orchestration technologies such as Docker and Kubernetes.
  • Familiarity with AI frameworks and libraries such as TensorFlow and PyTorch.
  • Knowledge of network and storage solutions that support AI and HPC workloads.
  • Familiarity with workload scheduling tools such as Slurm.

Tools and Technologies:

  • NVIDIA Mission Control / Bright Cluster Manager
  • Run:ai
  • Linux: Ubuntu and Red Hat Enterprise Linux (RHEL)
  • High-performance computing (HPC) platforms
  • Cloud and on-premises infrastructure
  • Docker and Kubernetes
  • TensorFlow and PyTorch
  • Slurm
  • AI compute, network, and storage infrastructure

Similar jobs