Quick Overview
Job Description
Platform Engineer - DevOps Specialist.
Job Type: Contract
Location: San Francisco, CA (Remote
End client : Pinterest
Role Summary:
We are seeking an experienced AI Infrastructure Engineer to support and optimize large-scale AI/ML platforms within AWS environments. This role focuses on AI model training and inference infrastructure, distributed training workloads, cloud-native platform engineering, and AWS AI accelerator technologies including Trainium and Inferentia. The ideal candidate will have experience building scalable AI infrastructure and supporting modern machine learning workloads at scale.
Key Responsibilities:
· Design, deploy, and support AI/ML infrastructure within AWS environments.
· Build and maintain scalable platforms for large-scale model training and inference workloads.
· Support distributed training environments for LLMs and Generative AI workloads.
· Implement and manage cloud-native infrastructure using Kubernetes, EKS, and containerized technologies.
· Utilize AWS Trainium, Inferentia, EC2 Trn instances, and AWS Neuron SDK to support AI workloads.
· Optimize AI platform performance through benchmarking, tuning, and infrastructure improvements.
· Support migration and optimization efforts involving AWS AI accelerator technologies.
· Collaborate with engineering teams to deliver scalable and reliable AI infrastructure solutions.
Required Qualifications:
· 7+ years of experience in Cloud, Data, or AI Infrastructure.
· Experience supporting large-scale AI/ML training workloads within AWS environments.
· Experience with AWS Trainium, Inferentia, EC2 Trn instances, and AWS Neuron SDK for AI model training and inference.
· Strong understanding of LLMs, Generative AI, distributed training, and AI/ML infrastructure.
· Hands-on experience with Kubernetes (EKS), Docker, Python, PyTorch, and cloud-native architectures.
· Experience with AI performance optimization, benchmarking, and infrastructure scalability.
· Exposure to AWS Trainium or Inferentia environments.
· Experience optimizing or migrating AI workloads to AWS AI accelerator platforms.
· Experience supporting large-scale distributed AI training environments.
Similar jobs
- ST
Remote Senior Databricks Platform Engineer-W2
NewSR Talent Solution Inc
United States🇺🇸Remote20 hours agoSQLAWSETL+7Technology - NI
C++ Engineer with SRE
NewNityo Infotech Corporation
San Jose, CA🇺🇸Hybrid20 hours agoMicroservicesSQLLoad Balancing+9Technology - KT
Director of DevOps
NewKforce Technology Staffing
Mesa, AZ🇺🇸Hybrid20 hours agoDockerSpinnakerAWS+16Technology - TS
Project Manager with Salesforce & Azure DevOps
NewTECHAGI SOLUTIONS LLC
Minneapolis, MN🇺🇸$60/hrHybrid20 hours agoScrumAgileAzure+1Technology - GL
DevOps/System Engineer - San Jose, CA - W2 - 3Months of Contract
NewGeopaq Logic
San Jose, CA🇺🇸Hybrid20 hours agoShellGitHTTP+2Technology - RI
DevOps Engineer
Raas Infotek LLC
Boston, MA🇺🇸On-site2 months agoDockerMicroservicesShell+23Technology - SN
Platform Engineer (Kong)
NewShree Narayani Networking Solutions LLC
United States🇺🇸$70/hrRemote20 hours agoDockerAPI GatewayAWS+15Technology - OK
Staff Site Reliability Engineer, Networking w/ active TS/SCI with Security Clearance
NewOkta, Inc.
Southern Md Facility, MD🇺🇸$174k - $238k/yrOn-site20 hours agoDockerAWSLoad Balancing+11Technology - PR
Oracle Database Platform Engineer - (Local to Massachusetts)
NewProhires
Boston, MA🇺🇸On-site20 hours agoOracleSQLShell+4Technology - MM
Senior DevOps Engineer
NewMitchell Martin, Inc.
United States🇺🇸$160k/yrRemote20 hours agoTechnology - AT
Azure Devops Engineer
NewAcadia Technologies, Inc.
Maryland City, MD🇺🇸Hybrid20 hours agoAzureTechnology - DI
Only W2- Cloud DevOps Engineer – Alibaba Cloud & AWS
NewDigitive LLC
Austin, TX🇺🇸On-site20 hours agoDockerShellAWS+10Technology