Haystack
← Back to Jobs
Remote
Technology
SA

Part time: Direct client requirement for Senior DevOps Engineer / DevOps Architect for Remote

SaksoftUnited States🇺🇸United StatesPosted 2 Sept 2026

Why This Role Stands Out

This remote, part-time role offers a unique opportunity to shape the future of eCommerce and AI-enabled DevOps, leveraging your extensive architectural expertise to drive innovation. You'll thrive here if you possess deep experience in cloud platforms, CI/CD, and cutting-edge AI technologies, allowing you to make a significant impact on advanced projects.

Quick Overview

Seniority
Mid Senior
Work mode
Remote
Location
United States
Posted
22 hours ago
DockerDynamoDBMicroservicesNode.jsSQLShellAPI GatewayAWSELKLoad BalancingNew RelicSplunkTCP/IPAzureBashCDNCloudFormationDNSDatadogGenerative AIGitGitHub ActionsGitLab CIGoogle CloudGrafanaHTTPHTTPSHelmJavaJenkinsKubernetesLLMPrometheusPythonRedisTerraformWAF

Job Description

Senior DevOps Engineer / DevOps Architect

eCommerce & AI-Enabled DevOps

Experience: 15+ Years  |  Domain: eCommerce / Digital Commerce  |  Role: Senior Technical / Architecture

 Parttime: 50 hours/Month 

Role Overview

We are looking for a highly experienced Senior DevOps Engineer / DevOps Architect with 15+ years of experience in enterprise eCommerce and digital commerce environments. The candidate should have strong hands-on and architectural expertise across AWS, Google Cloud Platform, Docker, Kubernetes, Infrastructure, CI/CD, Infrastructure as Code, Cloud Security, Monitoring, and Production Operations. The ideal candidate should also have practical experience in AI-enabled DevOps, including Generative AI, AI agents, AIOps, intelligent automation, predictive monitoring, automated incident analysis, and AI-assisted software delivery.

 

1. DevOps & Cloud Architecture

·       Define and implement scalable, secure, highly available DevOps and Cloud architectures for enterprise eCommerce platforms.

·       Design and manage cloud environments across AWS and Google Cloud Platform.

·       Define architecture for Development, QA, UAT, Pre-Production, and Production environments.

·       Design solutions covering compute, storage, networking, load balancing, CDN, DNS, databases, caching, security, monitoring, and disaster recovery.

·       Drive cloud modernization and migration initiatives.

·       Perform capacity planning and infrastructure optimization for high-volume eCommerce platforms.

 

2. AWS & Google Cloud Platform

·       Strong hands-on experience with AWS services including EC2, ECS/EKS, Lambda, S3, RDS, DynamoDB, ElastiCache, CloudFront, Route 53, API Gateway, CloudWatch, IAM, VPC, ALB/NLB, WAF, and Secrets Manager.

·       Experience with Google Cloud Platform services including GKE, Compute Engine, Cloud Storage, Cloud SQL, Cloud Load Balancing, Cloud CDN, Cloud Monitoring, IAM, VPC, and Secret Manager.

·       Experience with multi-cloud architecture and workload optimization is preferred.

 

3. Kubernetes & Docker

·       Design, deploy, and manage enterprise Kubernetes environments.

·       Strong experience with AWS EKS and Google Cloud Platform GKE.

·       Develop and manage Docker containers and containerized applications.

·       Work with Pods, Deployments, Services, Ingress, ConfigMaps, Secrets, StatefulSets, Helm, HPA, and Persistent Volumes.

·       Implement container security and image management.

·       Troubleshoot Kubernetes cluster, application, networking, and performance issues.

·       Implement autoscaling and resilient container architectures.

 

4. Infrastructure & Networking

·       Strong understanding of Linux/Unix administration, web/application servers, database infrastructure, storage, compute, virtualization, load balancers, high availability, disaster recovery, backup/restoration, and capacity planning.

·       Strong networking knowledge covering TCP/IP, DNS, HTTP/HTTPS, SSL/TLS, VPN, VPC/VNet, subnets, routing, NAT, firewalls, security groups, WAF, CDN, proxy, and load balancing.

·       Ability to troubleshoot issues across application, container, cloud, infrastructure, and network layers.

 

5. CI/CD & Release Automation

·       Design and implement enterprise CI/CD pipelines.

·       Automate build, testing, security scanning, deployment, rollback, and release processes.

·       Experience with Jenkins, GitHub Actions, GitLab CI/CD, Azure DevOps, and Argo CD.

·       Implement Blue/Green, Canary, Rolling, and zero-downtime deployments.

·       Integrate automated testing and security gates into pipelines.

·       Improve deployment frequency, release quality, and engineering productivity.

 

6. Infrastructure as Code

·       Strong experience with Terraform and/or CloudFormation.

·       Create reusable infrastructure modules and automate cloud provisioning.

·       Manage infrastructure through Git-based version control.

·       Establish enterprise IaC standards and governance.

·       Automate environment creation and configuration.

 

7. AI-Driven DevOps / AIOps

·       Use Generative AI tools to accelerate infrastructure scripting, Terraform development, Kubernetes configuration, CI/CD pipeline creation, shell/Python scripting, troubleshooting, and documentation.

·       Design or implement AI agents for automated log analysis, incident investigation, Root Cause Analysis, infrastructure health checks, deployment validation, CI/CD failure analysis, automated remediation, performance analysis, security issue identification, environment validation, and release readiness checks.

·       Implement AIOps capabilities including predictive monitoring, anomaly detection, intelligent alert correlation, incident prediction, event correlation, automated RCA, capacity forecasting, performance optimization, intelligent ticket classification, and automated incident response.

 

8. AI-Enabled CI/CD

·       Drive intelligent CI/CD pipelines using AI.

·       Implement workflows such as Code Commit → AI Code Analysis → Build → Automated Testing → Security Scan → AI Risk Analysis → Deployment → AI Deployment Validation → Monitoring → Automated Remediation.

·       Use AI to identify high-risk deployments, potential production issues, failed deployment patterns, infrastructure configuration issues, performance degradation, and security vulnerabilities.

 

9. AI-Based Production Support

·       Build AI-driven production monitoring and support capabilities.

·       Implement intelligent analysis of application and infrastructure logs.

·       Correlate logs, metrics, traces, alerts, deployment history, and infrastructure changes using AI.

·       Develop automated RCA and incident-summary generation.

·       Build self-healing capabilities where appropriate.

·       Reduce MTTR through intelligent incident diagnosis and automation.

 

10. eCommerce DevOps Expertise

·       Strong understanding of DevOps requirements for enterprise eCommerce platforms such as Salesforce Commerce Cloud, HCL Commerce, Adobe Commerce/Magento, SAP Commerce/Hybris, Medusa Commerce, and custom Java/Node.js commerce platforms.

·       Experience with product catalog, search, inventory, order management, payment systems, API integrations, microservices, Redis/cache, Elasticsearch/Solr, CDN, databases, batch processing, and third-party integrations.

·       Understand peak traffic, flash sales, promotional events, Black Friday/holiday traffic, performance testing, autoscaling, and high availability.

 

11. Monitoring & Observability

·       Implement enterprise observability using Prometheus, Grafana, ELK/Elastic, Splunk, Datadog, New Relic, CloudWatch, and Cloud Monitoring.

·       Implement infrastructure, application, Kubernetes, distributed tracing, log aggregation, alert management, SLI/SLO monitoring, and performance dashboards.

·       Leverage AI for intelligent alerting, anomaly detection, and predictive monitoring.

 

12. Security & DevSecOps

·       Implement DevSecOps practices across the SDLC.

·       Integrate security scanning into CI/CD.

·       Secure containers and Kubernetes clusters.

·       Implement IAM and least-privilege access.

·       Manage secrets and certificates securely.

·       Implement vulnerability scanning and remediation.

·       Support cloud security and compliance requirements.

 

13. Production & Incident Management

·       Provide technical leadership during P1/P2 production incidents.

·       Perform detailed Root Cause Analysis and define permanent corrective actions.

·       Support production releases and hypercare.

·       Develop operational runbooks.

·       Implement disaster recovery and business continuity processes.

·       Drive reduction in MTTR, deployment failures, and production incidents.

 

14. Technical Leadership

·       Lead and mentor DevOps and Infrastructure teams.

·       Define enterprise DevOps standards and best practices.

·       Conduct architecture and technical design reviews.

·       Partner with Application, QA, Security, Architecture, and Delivery teams.

·       Lead cloud migration and modernization initiatives.

·       Support technical estimations, proposals, RFP/RFQ responses, and pre-sales activities.

·       Define and drive AI-enabled DevOps transformation initiatives.

 

Required Technical Skills

Area

Skills

Experience

15+ Years

Domain

eCommerce / Digital Commerce

Cloud

AWS, Google Cloud Platform

Containers

Docker, Kubernetes

Kubernetes

EKS, GKE, Helm, Ingress, HPA

IaC

Terraform, CloudFormation

CI/CD

Jenkins, GitHub Actions, GitLab, Argo CD

OS

Linux / Unix

Scripting

Shell/Bash, Python

Networking

TCP/IP, DNS, VPN, SSL, CDN, WAF, Load Balancers

Monitoring

Prometheus, Grafana, ELK, Datadog, New Relic

Security

IAM, DevSecOps, Secrets, Vulnerability Management

Version Control

Git, GitHub, GitLab, Bitbucket

AI / GenAI

Generative AI, AI Agents, AIOps, LLM-based automation

Automation

Infrastructure & Deployment Automation

eCommerce

Enterprise Commerce Platforms

 

AI / DevOps Skill Expectations

·       GenAI for DevOps: AI-assisted scripting, AI-generated Terraform/Kubernetes configurations, AI-assisted troubleshooting, and AI documentation.

·       AI Agents: Incident Agent, Monitoring Agent, Deployment Agent, RCA Agent, Infrastructure Health Agent, and Release Validation Agent.

·       AIOps: Anomaly Detection, Predictive Monitoring, Intelligent Alert Correlation, Automated RCA, and Auto-remediation.

·       AI-Enabled SDLC: AI-based Code Review, Test Automation, Security Analysis, Deployment Risk Assessment, and Production Monitoring.

 

Key Success Metrics

·       99.9%+ platform availability.

·       Reduction in deployment failures and production incidents.

·       Reduction in MTTR.

·       Increased deployment frequency and automation coverage.

·       Reduced manual infrastructure activities.

·       Improved cloud infrastructure utilization and cost optimization.

·       AI-driven reduction in DevOps operational effort.

·       Faster incident identification and resolution.

·       Improved production stability and scalability.

 

Ideal Candidate Profile

 

15+ Years | Enterprise eCommerce | AWS | Google Cloud Platform | Docker | Kubernetes | EKS/GKE | Terraform | CI/CD | Infrastructure | Networking | DevSecOps | Observability | AIOps | GenAI | AI Agents | Production Support | Cloud Architecture

Positioning: This is a Senior DevOps / DevOps Architect role with strong infrastructure ownership and AI transformation capability, rather than a conventional DevOps Engineer focused primarily on CI/CD.

 

Similar jobs