Quick Overview
Job Description
Position: AI Infrastructure Engineer
Location: Remote
Hiring Mode: 6+ Months Contract
Job Description:
The AI Infrastructure Operations Engineer will manage, maintain, and optimize artificial intelligence (AI) infrastructure across on-premises and cloud environments. This hands-on role focuses on platform availability, performance, resource utilization, troubleshooting, operational excellence, and team enablement. The engineer will support AI workloads and model deployments while providing practical guidance and training on infrastructure tools and best practices. This position emphasizes reliable operations and knowledge transfer rather than strategic development or workload definition.
Responsibilities:
- Manage and maintain AI infrastructure to support high availability, reliability, and performance.
- Implement, administer, and optimize AI operations using NVIDIA Mission Control, Bright Cluster Manager, Run:ai, and related platform tools.
- Collaborate with cross-functional teams to support AI workloads and improve compute, accelerator, storage, and scheduling resource utilization.
- Provide training, mentoring, and hands-on guidance to team members on AI infrastructure tools, operational practices, and troubleshooting methods.
- Monitor platform health and system performance, diagnose issues, and resolve incidents to minimize downtime and improve resource allocation.
- Assist with the deployment, operation, and scaling of AI models and applications across supported environments.
- Evaluate relevant advancements in AI infrastructure technology and recommend practical operational improvements.
- Develop and maintain clear documentation for configurations, procedures, troubleshooting, and infrastructure management best practices.
Required Qualifications:
- Proven experience managing AI infrastructure and day-to-day platform operations.
- Hands-on proficiency with NVIDIA Mission Control or Bright Cluster Manager and Run:ai.
- Proficiency administering Linux operating systems, including Ubuntu and Red Hat Enterprise Linux (RHEL).
- Strong understanding of high-performance computing (HPC) environments and resource management concepts.
- Experience supporting both cloud platforms and on-premises infrastructure.
- Strong troubleshooting, analytical, and problem-solving skills with close attention to detail.
- Ability to collaborate effectively across technical teams and communicate clearly with varied audiences.
- Experience training and mentoring technical team members.
- Bachelor s degree in Computer Science, Engineering, or a related field, or equivalent professional experience.
Preferred Qualifications:
- Experience with containerization and orchestration technologies such as Docker and Kubernetes.
- Familiarity with AI frameworks and libraries such as TensorFlow and PyTorch.
- Knowledge of network and storage solutions that support AI and HPC workloads.
- Familiarity with workload scheduling tools such as Slurm.
Tools and Technologies:
- NVIDIA Mission Control / Bright Cluster Manager
- Run:ai
- Linux: Ubuntu and Red Hat Enterprise Linux (RHEL)
- High-performance computing (HPC) platforms
- Cloud and on-premises infrastructure
- Docker and Kubernetes
- TensorFlow and PyTorch
- Slurm
- AI compute, network, and storage infrastructure
Similar jobs
- TE
Gen AI Engineer
NewTechVirtue LLC
Fort Worth, TX🇺🇸Hybrid21 hours agoMicroservicesAWSMachine Learning+12Technology - MC
AI Engineer
NewMcKinsol Consulting Inc
Atlanta, GA🇺🇸On-site21 hours agoGenerative AIJavaScriptLLM+2Technology - CS
AI engineer and COACH with Security Clearance
NewCEdge Software Consultants
Saint Louis, MO🇺🇸Hybrid21 hours agoC++JavaTechnology - KS
Azure Databricks & Agentic AI Architect
NewKeypixel Software Solutions
Chicago, IL🇺🇸On-site21 hours agoSQLETLMLOps+5Technology - TA
AI Prompt Engineer Agentic AI Engineer
NewThe Avian Consulting LLC
Irving, TX🇺🇸Hybrid21 hours agoAWSAzureGoogle Cloud+2Technology - 3A
Java AI Developer
New3AMIGOSIT LLC
Charlotte, NC🇺🇸Hybrid21 hours agoMicroservicesMongoDBOracle+10Technology