Haystack
← Back to Jobs
Remote
Technology

AI/ML Platform Engineer - Brazil (Remote)

Georgia ITUnited States🇺🇸United StatesPosted 14 Aug 2026

Quick Overview

Work Type
Remote
Level
Mid Senior

Job Description

AI/ML Platform Engineer

Location: Brazil
Employment Type: Contract
Work Model: Remote

Role Overview

We are seeking an experienced AI/ML Platform Engineer to design, build, and maintain scalable AI/ML infrastructure and delivery platforms. The ideal candidate will have hands-on experience supporting GPU workloads, LLM integrations, Kubernetes, Docker, Helm, AWS, Terraform/Ansible, CI/CD, and DevOps tooling.

Key Responsibilities

  • Design and engineer scalable infrastructure for AI/ML workloads and GPU-based environments.
  • Build and maintain platforms supporting LLM applications and integrations.
  • Deploy and manage AI/ML workloads using Kubernetes, Docker, and Helm.
  • Develop and maintain CI/CD pipelines for ML models, applications, and infrastructure.
  • Automate infrastructure provisioning and configuration using Terraform and Ansible.
  • Deploy and manage AI/ML infrastructure on AWS.
  • Implement monitoring, logging, security, and performance optimization for AI/ML platforms.
  • Support model deployment, integration, scaling, and operationalization.
  • Troubleshoot infrastructure, container, Kubernetes, and deployment issues.
  • Establish DevOps and MLOps best practices for reliable AI/ML delivery.
  • Collaborate with data scientists, ML engineers, software engineers, and DevOps teams.

Required Skills

  • Strong experience as an AI/ML Platform Engineer, MLOps Engineer, or DevOps Engineer supporting AI/ML platforms.
  • Hands-on experience with:
    • Kubernetes
    • Docker
    • Helm
    • AWS
    • Terraform and/or Ansible
    • CI/CD
    • GPU workloads
    • LLM integrations
  • Strong understanding of cloud infrastructure, containers, orchestration, and automation.
  • Experience building scalable and highly available AI/ML platforms.
  • Strong scripting and automation skills.
  • Knowledge of monitoring, logging, security, and infrastructure optimization.

Preferred Skills

  • Experience with MLOps platforms and practices.
  • Experience with NVIDIA GPU infrastructure and CUDA environments.
  • Knowledge of Python and AI/ML frameworks.
  • Experience with LLM platforms, model serving, and inference workloads.
  • Familiarity with AWS services such as EKS, EC2, S3, and SageMaker.
  • Experience with Prometheus, Grafana, or similar observability tools.

Skills

Docker
AWS
MLOps
Ansible
CUDA
Grafana
Helm
Kubernetes
LLM
Prometheus
Python
Terraform

Similar jobs