Haystack
← Back to Jobs
Technology
AT

Platform Engineer - DevOps Specialist

Aivanta Tech IncUnited States🇺🇸United StatesPosted Oct 9, 2026

Quick Overview

Seniority
Mid Senior
Work mode
Hybrid
Location
United States
Posted
20 hours ago
DockerAWSMachine LearningGenerative AIKubernetesPyTorchPython

Job Description

Platform Engineer - DevOps Specialist.

Job Type: Contract

Location: San Francisco, CA (Remote

End client : Pinterest

Role Summary:
We are seeking an experienced AI Infrastructure Engineer to support and optimize large-scale AI/ML platforms within AWS environments. This role focuses on AI model training and inference infrastructure, distributed training workloads, cloud-native platform engineering, and AWS AI accelerator technologies including Trainium and Inferentia. The ideal candidate will have experience building scalable AI infrastructure and supporting modern machine learning workloads at scale.

 

Key Responsibilities:

·  Design, deploy, and support AI/ML infrastructure within AWS environments.

·  Build and maintain scalable platforms for large-scale model training and inference workloads.

·  Support distributed training environments for LLMs and Generative AI workloads.

·  Implement and manage cloud-native infrastructure using Kubernetes, EKS, and containerized technologies.

·  Utilize AWS Trainium, Inferentia, EC2 Trn instances, and AWS Neuron SDK to support AI workloads.

·  Optimize AI platform performance through benchmarking, tuning, and infrastructure improvements.

·  Support migration and optimization efforts involving AWS AI accelerator technologies.

·  Collaborate with engineering teams to deliver scalable and reliable AI infrastructure solutions.

 

Required Qualifications:

·  7+ years of experience in Cloud, Data, or AI Infrastructure.

·  Experience supporting large-scale AI/ML training workloads within AWS environments.

·  Experience with AWS Trainium, Inferentia, EC2 Trn instances, and AWS Neuron SDK for AI model training and inference.

·  Strong understanding of LLMs, Generative AI, distributed training, and AI/ML infrastructure.

·  Hands-on experience with Kubernetes (EKS), Docker, Python, PyTorch, and cloud-native architectures.

·  Experience with AI performance optimization, benchmarking, and infrastructure scalability.

·  Exposure to AWS Trainium or Inferentia environments.

·  Experience optimizing or migrating AI workloads to AWS AI accelerator platforms.

·  Experience supporting large-scale distributed AI training environments.

Similar jobs