Haystack
← Back to Jobs
Technology
DT

Platform Engineer / SRE (Site Reliability Engineer) Cloud & AI Automation

Donato Technologies IncDenver, CO🇺🇸United StatesPosted 24 Aug 2026

Quick Overview

Seniority
Mid Senior
Work mode
Hybrid
Location
Denver, CO, United States
Posted
17 hours ago
DockerMicroservicesAWSELKETLEncryptionSOC 2SnowflakeAirflowAzureBigQueryCloudFormationDatadogGDPRGitHub ActionsGitLab CIGoogle CloudGrafanaHIPAAJenkinsKibanaKubernetesPrometheusPythonRedshiftTerraform

Job Description

Platform Engineer / SRE (Site Reliability Engineer) – Cloud & AI Automation 

Location: Colorado (Mon to Thu wfo)

Role Overview:

  • We are seeking a highly skilled and proactive **Senior Platform Engineer / SRE** to lead the design, automation, and operational excellence of cloud platforms and AI-driven systems.
  • The ideal candidate will be a seasoned engineer with hands-on experience in **AWS, Terraform, Kubernetes, CI/CD, Infrastructure as Code (IaC), and Observability**, who thrives in fast-paced, high-impact environments.
  • You will play a pivotal role in building and maintaining scalable, secure, and self-healing platforms that support enterprise-grade applications and AI-powered workflows.
  • This role sits at the intersection of **platform engineering, reliability, security, and AI automation**, where you’ll drive innovation through agentic AI, infrastructure automation, and cloud cost optimization.

Key Responsibilities:

- Design, implement, and manage **cloud-native platforms** on AWS using **Terraform, CloudFormation, and IaC best practices**.

- Lead **infrastructure automation** initiatives to enable zero-touch provisioning, self-service environments, and CI/CD pipeline integration.

- Architect and operate **highly available, scalable, and secure systems** using **Kubernetes (EKS/ECS), Docker, and microservices**.

- Implement and maintain **robust observability stacks** using **Grafana, Datadog, Prometheus, ELK, and OpenTelemetry** for real-time monitoring, alerting, and performance tuning.

- Own **incident management lifecycle**: lead on-call rotations, conduct **RCA (Root Cause Analysis)**, and implement preventive measures to improve system reliability.

- Drive **cloud cost optimization** strategies across multi-cloud environments (AWS, Azure, Google Cloud Platform) through right-sizing, auto-scaling, tagging policies, and usage analytics.

- Enforce **security compliance standards** including **PCI, PII, HIPAA, GDPR, ISO 27001/27701**, and SOC 2, ensuring IAM policies, encryption, and audit readiness.

- Integrate **Okta, AWS IAM, and identity federation** for secure access control and role-based access management.

- Spearhead **AI automation and agentic AI** use cases for infrastructure operations—automating deployments, incident response, policy enforcement, and resource provisioning.

- Collaborate with DevOps, SRE, Security, and Product teams to deliver **platform capabilities** that accelerate time-to-market and improve developer experience.

- Mentor junior engineers and contribute to **engineering best practices, documentation, and knowledge sharing**.

Required Qualifications:

- Bachelor’s or Master’s degree in Computer Science, Engineering, or related field.

- 6+ years of hands-on experience in **DevOps, SRE, or Platform Engineering** roles.

- Expertise in **AWS cloud services** (EC2, S3, Lambda, RDS, VPC, IAM, CloudFront, etc.) and **Terraform/CloudFormation** for infrastructure provisioning.

- Proven experience with **Kubernetes (EKS, AKS, GKE)**, **Docker**, and container orchestration.

- Strong proficiency in **CI/CD pipelines** using Jenkins, GitHub Actions, GitLab CI, or similar tools.

- Deep understanding of **observability tools**: Grafana, Datadog, Prometheus, Kibana, and logging frameworks.

- Experience with **incident management, RCA, and post-mortem processes** in production environments.

- Solid knowledge of **security compliance frameworks**: PCI-DSS, HIPAA, GDPR, CCPA, ISO 27001/27701.

- Experience integrating **identity providers (Okta, AWS IAM)** and managing RBAC across cloud environments.

- Hands-on experience with **Python** for automation, scripting, and tooling.

- Familiarity with **Agentic AI, AI automation, and intelligent operations (AIOps)** for infrastructure and platform management.

- Strong communication, collaboration, and leadership skills.

Preferred Qualifications:

- Experience with **multi-cloud environments** (AWS, Azure, Google Cloud Platform).

- Knowledge of **serverless architectures** (AWS Lambda, Azure Functions).

- Experience with **data platforms** (Snowflake, Redshift, BigQuery) and ETL pipelines (Airflow, Glue).

- Exposure to **AI/ML platforms** and **GenAI integration** in DevOps workflows.

- Certification in AWS, Google Cloud, or Kubernetes (e.g., AWS Certified DevOps Engineer, CKAD, CKA).

Similar jobs