Quick Overview
Job Description
Job Title: Senior Automation Engineer ( CyberArk , Red Hat, OpenShift , Azure)
Location: Remote
Duration: long term
About the Role
The Senior Automation Engineer is a hands-on technical delivery role on Client OpenShift Platform Readiness Program for Dominion Energy, a 32-week engagement to establish Red Hat OpenShift as Dominion’s enterprise application hosting platform across on-premises, Azure, and AWS. The program runs two parallel delivery tracks: Track 1 stabilizes the existing on-premises OpenShift environment and takes it through an Operational Readiness Review (ORR) to production go-live for Electric Transmission (ET) workloads; Track 2 delivers production-ready ARO and ROSA platforms through repeatable, automated deployment patterns. In this role you build the automation artifacts that underpin the platform operating model across all three environments.
Key Responsibilities
GitOps & Multi-Environment Delivery
· Refactor existing GitOps repositories in Azure DevOps to the approved branching model, environment promotion workflow, and merge approval gates
· Implement ArgoCD ApplicationSets using cluster generators and environment labels to target on-premises, ARO, and ROSA clusters from a common repository structure
· Implement the app-of-apps pattern for platform baseline configuration management
· Update ArgoCD RBAC to align with the agreed ownership model across the platform team, COE, and application teams
· Write Git repository documentation, including README files, branching strategy guide, contribution guide, and environment promotion runbooks
Namespace Onboarding & Platform Guardrails
· Build a standardized, automated namespace onboarding pipeline that delivers baseline RBAC, NetworkPolicies, resource quotas, and storage quotas consistently across all three environments
· Develop RBAC templates and NetworkPolicy configurations aligned to the agreed network policy ownership model across platform, COE, and application teams
· Implement Kyverno governance policies for platform-level guardrails, such as resource quota enforcement, image registry restrictions, and label and annotation standards
· Implement DomSI CA injection into pod trust stores via Kyverno (or the approved equivalent), enabling removal of the Broadcom SSL inspection bypasses
· Structure onboarding pipeline code, configuration, and documentation in a format compatible with Red Hat Developer Hub catalog integration
· Contribute to platform standards documentation covering image registry usage, network policy expectations, resource quota guidance, and GitOps promotion requirements
· Secrets, Certificates & Observability Integration
· Implement the CyberArk Conjur CSI secretless pattern as the primary secrets management approach, and ESO with Reloader for exception use cases, on-premises and on ARO and ROSA
· Configure cert-manager integrated with Dominion’s enterprise PKI, including automated renewal workflows and expiry alerting; support KUBE+ integration for front-end certificate management
· Automate Dynatrace Operator deployment and OneAgent rollout, and deliver dashboard and alerting configurations as code
· Support Splunk forwarding for platform audit events, compliance findings, and ACS runtime alerts
· Support Compliance Operator profile configurations for NIST scanning and ACS policy configurations as version-controlled artifacts
Automation Artifact Library & Platform Readiness
· Maintain version control hygiene across the Azure DevOps automation artifact library, including pull request reviews, README documentation, and contribution guide maintenance
· Develop Ansible Automation Platform playbooks and Terraform modules as reusable artifacts for Day-2 operations and provisioning workflows, with scope and format confirmed during Planning and Design
· Validate automation patterns in the client Advanced Technology Center (ATC) before they are introduced into Dominion’s production environments
· Support NetApp Trident StorageClass and snapshot policy definitions and Rubrik KUPR backup integration across on-premises, ARO, and ROSA
· Support the reference application deployment, the ARO availability-zone failover demonstration, and ORR evidence for ET production go-live
· Support namespace provisioning and onboarding as non-ET applications migrate from bare metal to virtual clusters and during hypercare
· Governance & Delivery
· Participate fully in Agile ceremonies (sprint planning, sprint reviews and demos, retrospectives, and backlog refinement) on a two-week cadence aligned to Dominion’s sprint cycle
· Work from the project backlog: pull stories that meet the Definition of Ready, and close them only when they meet the Definition of Done, including testing evidence
· Demonstrate working, tested increments at every sprint close; your sprint artifacts feed velocity and burndown metrics, the weekly status report, and Dominion’s program milestone framewok
· Flag risks, blockers, and dependencies early, including infrastructure readiness gates for VCF, NetApp, and network/firewall, and escalate through the Delivery Lead rather than absorbing silently
· Route scope changes through the agreed change control process, and operate within Dominion’s change management practices for production work
Required Technical Skills
· Automation & Orchestration
· Skill Area
· Expected Depth
· Ansible / AAP (strong) – 5+ yrs
· Production-grade playbook and role authoring. Comfortable with block/rescue/always, argument specs, handlers, and Ansible Vault. Fluent with Ansible Automation Platform (AAP) job templates, credentials, and inventories, including exposing templates through the AAP API so pipelines can call them for Day-2 cluster and OS operations.
· GitOps with ArgoCD – 3+ yrs
· Multi-cluster deployments using ApplicationSets (cluster generators, environment labels), App-of-Apps, sync waves, ArgoCD RBAC, and drift reconciliation. Experience with canary and blue/green rollout strategies.
· Terraform – 3+ yrs
· AWS and Azure networking and cluster-adjacent infrastructure. Remote state, workspace isolation, and versioned modules shared across teams.
· CI/CD pipelines – 5+ yrs
· Authoring and maintaining pipelines (Azure DevOps Pipelines preferred; GitHub Actions, GitLab CI, or Jenkins also relevant) for container build, test, and publish, feeding multi-environment Helm and Kustomize deployments. Understanding of artifact promotion, deployment gates, and environment parity.
· Git / Azure DevOps Repos – 5+ yrs
· Branching strategies (trunk-based or GitFlow), pull requests and merge approval gates, code review practice, rebase vs. merge, conflict resolution. Comfortable with pre-commit hooks.
· Python or Bash – 5+ yrs
· Python 3.10+ or Bash for automation glue, API integrations, custom filter plugins, and test tooling. Comfortable with requests, click/argparse, and pydantic.
· Testing frameworks – 5+ yrs
· Molecule for Ansible roles, pytest for Python. Able to write tests from acceptance criteria, not just to cover implementation.
· Container Platforms
· Platform
· Expected Depth
· OpenShift 4.x / Kubernetes – 5+ yrs
· Install, upgrade, and operate production clusters on VMware VCF/vSphere, ARO, and ROSA. Operator lifecycle management, MachineSets and MachinePools, node lifecycle, and capacity planning. Red Hat Advanced Cluster Management (ACM) placement policies a plus.
· Security, Policy & Multi-Tenancy
· RBAC, namespaces and projects, ServiceAccounts, and NetworkPolicies that let developers self-serve inside defined guardrails. Kyverno policy authoring; working knowledge of Red Hat ACS and the OpenShift Compliance Operator.
· Secrets Management
· CyberArk Conjur (CSI secretless pattern) and External Secrets Operator (ESO) with Reloader, or comparable secrets-injection patterns.
· Ingress, DNS & Certificates
· OpenShift-native ingress (Router) and NGINX; cert-manager with ACME or enterprise PKI; CA injection into pod trust stores.
· Storage & DR
· StorageClasses, CSI drivers (NetApp Trident preferred), PV/PVC management, snapshots, and backup and restore with Rubrik KUPR or Velero/OADP.
· GPU Workloads (nice-to-have)
· NVIDIA GPU Operator, including MIG and time-slicing configurations.
Source of Truth & Tooling
· Azure DevOps as the source of truth; GitOps repositories, merge approval gates, and the version-controlled automation artifact library
· ServiceNow integration patterns for change records, incident creation, and CMDB reconciliation
· Helm and Kustomize for application packaging and environment-specific overlays
· Observability: Dynatrace (Operator, OneAgent, dashboards, and alerting) preferred, with Datadog or Grafana experience transferable; Splunk for audit and compliance event forwarding; SLO / error-budget tracking
· DevOps Practices
· Secrets management: CyberArk Conjur, HashiCorp Vault, or equivalent; never plaintext credentials in code, inventory, or tickets
· Observability: structured logging, correlation IDs, and audit trails across Azure DevOps, Splunk, Jira, or equivalent
· Linting and code quality: ansible-lint (production profile), yamllint, black/ruff, pre-commit hooks enforced in CI
· Compliance hardening: applying NIST controls (including via the OpenShift Compliance Operator) and DISA STIG to clusters, nodes, and pipelines
· Experience & Qualifications
· Minimum Experience
· 10+ years of infrastructure or platform engineering experience, with at least 5 years running OpenShift or Kubernetes in production
· Demonstrated experience taking automation code from lab/proof-of-concept to production operation at scale
· Hands-on experience operating Kubernetes at scale, across many clusters and regions
· Prior experience operating in an Agile / Scrum environment as a contributing engineer, not just as a participant
· Experience working in regulated industries (utilities and critical infrastructure, financial services, healthcare, public sector, or similar) with frameworks such as PCI and NIST 800-171, or comparable change-control discipline
· Preferred Qualifications
· Red Hat Certified Specialist in OpenShift Administration (EX280)
· Red Hat Certified Engineer / Red Hat Certified Specialist in Ansible Automation
· Experience with Azure Red Hat OpenShift (ARO), Red Hat OpenShift Service on AWS (ROSA), or Azure Government
· Experience running backup and disaster recovery programs for OpenShift with Rubrik KUPR or OADP/Velero
· Hands-on experience with Red Hat ACS, Kyverno, and the OpenShift Compliance Operator
· Experience migrating application workloads from bare metal to virtualized OpenShift clusters
· Exposure to Red Hat Developer Hub (Backstage) catalog patterns
· Exposure to event-driven automation patterns (EDA with Ansible Rulebooks, webhook-triggered remediation, or equivalent)
Similar jobs
- B/
Expert Systems Engineer with Security Clearance
NewB/Core
Herndon, VA🇺🇸$170k - $195k/yrHybrid18 hours agoAWSSplunkAnsible+4Technology - VE
Embedded Software Engineers
NewVertogic
North Reading, MA🇺🇸$60/hrHybrid18 hours agoEmbedded SystemsAgileC+++1Technology - SI
Automation Engineer
NewSilverSearch, Inc.
Jersey City, NJ🇺🇸Hybrid18 hours agoMicroservicesSQLAWS+12Technology - C&
RPG Programmer Analyst
NewCox-Little & Company
State College, PA🇺🇸HybridYesterdaySOAPREST - WY
System Engineer 1 with Security Clearance
NewWyetech, LLC
Columbia, MD🇺🇸$35 - $106/hrHybrid18 hours agoTechnology - KT
Senior Software Developer
NewKforce Technology Staffing
Davie, FL🇺🇸Hybrid18 hours agoMongoDBSQLSQL Server+11Technology