Quick Overview
Job Description
Employment Eligibility Statement:
Due to specific project and client requirements, this position is open to U.S. Citizens and U.S. Lawful Permanent Residents (s). Sponsorship is not available at this time.
Danta Technologies evaluates all candidates in compliance with the Immigration and Nationality Act (INA) and EEOC guidelines. All hiring decisions are made without regard to race, color, religion, sex, gender identity, sexual orientation, national origin, age, disability, veteran status, or any other protected characteristic.
Note:
This position is based in Bellevue, WA, Overland Park, KS, or Dallas , TX and requires 5 days per week working from the office.
This position is based in the USA, and the client is allowing candidates to work remotely within the USA
Only candidates who are currently located within the United States will be considered.
Applications from candidates residing outside the USA will not be accepted.
Client Expectations:
The client wants a Principal/Staff-level Observability & Platform Architect who has actually run VictoriaMetrics Cluster in production at scale, can architect the entire observability platform, and has enough Kubernetes/OpenShift, security, networking, GitOps, telecom, and architecture-governance experience to own the platform end-to-end.
Job Description :
| Expert | Deep, hands-on production ownership. Design and delivery-ready from day one. | Intermediate | Solid operational experience. Productive from day one; deepens in role. |
Victoria Metrics:
Technical Skills
Hands-on operation of Victoria Metrics in production — VM Cluster topology (VM Insert, VM Storage, VM Select), VM agent scrape and stream aggregation, VM Auth multi-tenant routing, VM Alert and VM Alert manager rule design, and Metrics QL query authoring. Ability to govern active time series cardinality, configure retention and downsampling, and design multi-cluster federation with cross-cluster deduplication and write-path failover.
Experience Profile:
Verifiable production VM Cluster deployments handling sustained, high-cardinality workloads — not evaluation or sandbox only. Evidence of diagnosing and remediating a cardinality problem in a live environment. Operational experience of VM Auth-based per-tenant data isolation beyond authentication alone. Familiarity with the Victoria Metrics Operator on Kubernetes and its upgrade lifecycle. Intermediate accepted where Expert-level Observability is demonstrated.
| Domain | Level | Technical Skills | Experience Profile |
| Observability | Expert | Expert-level design and operation of full-stack observability platforms: Prometheus scrape and federation, Grafana dashboard and datasource provisioning at scale, SLI/SLO definition and error-budget alerting, structured logging pipelines, and distributed tracing integration. Ability to define metric taxonomies and labelling standards that remain coherent across heterogeneous multi-vendor sources. |
Delivered observability platforms serving production engineering or operations teams in a large-scale environment. Designed SLO-based alerting frameworks that measurably reduced operational noise and accelerated incident response. Led metric standardisation across multiple teams or vendor sources. Evidence of instrumenting the observability platform itself — not only the workloads it serves. |
| Architecture Leadership | Expert | Ability to serve as the single accountable design authority for a complex technical platform — producing High-Level Designs, Low-Level Designs, Architecture Decision Records, sequence diagrams, and interface contracts that engineering teams build from directly. Comfortable presenting architecture decisions and their trade-offs to mixed audiences from senior engineers to executive sponsors, adjusting depth without losing accuracy. |
Has held a Principal or Staff Architect role where design and build were explicitly separated. Maintained an ADR record throughout a programme delivery. Can name a design decision they were challenged on, the trade-offs documented, and how the decision held up through delivery. Evidence of architecture governance participation: design review boards, security assurance gates, and formal approval before engineering investment began. |
| Security & Identity | Expert | Zero Trust architecture as a design default — every inter-component path authenticated, encrypted, and auditable. Deep integration with HashiCorp Vault including the Vault Secrets Operator synchronisation pattern, dynamic secrets, and PKI secrets engine for automated certificate issuance. Enterprise PKI lifecycle: certificate generation, renewal automation, CA distribution to distributed cluster workloads, and expiry alerting. OIDC and OAuth2 federation eliminating all static credential and local account patterns. |
Integrated Vault with Kubernetes workloads in production using the VSO sync pattern, not solely as a secrets store. Designed and operated automated TLS certificate lifecycle across a distributed platform including renewal without service interruption. Implemented OIDC federation for platform services against an enterprise identity provider. Evidence of security governance contribution: compliance reviews, threat model participation, or security assurance gate ownership. |
| Kubernetes & OpenShift | Intermediate | Kubernetes administration and troubleshooting across StatefulSets, PersistentVolumeClaims, StorageClasses, Operators, RBAC, and NetworkPolicy. Red Hat OpenShift-specific competency: Security Context Constraints, OCP upgrade path management, OLM operator lifecycle, and the material differences from cloud-managed Kubernetes that affect stateful, high-throughput workloads. Multi-cluster topology design including hub and spoke architectures and cross-cluster connectivity. |
Operational experience on Red Hat OpenShift in an on-premises or bare-metal enterprise environment — not exclusively cloud-managed Kubernetes. Has configured SCCs for a stateful workload without cluster-admin workarounds. Managed an OCP cluster through at least one version upgrade including Operator compatibility validation. Experience with OpenShift ACM or equivalent for policy and workload deployment across a cluster fleet. |
| GitOps & CI/CD | Intermediate | GitOps-native delivery model using ArgoCD or FluxCD as the reconciliation engine — all cluster state managed in Git, no manual production changes permitted. Kustomize overlay strategy for multi-environment and multi-tenant deployments: base resource definitions with environment-specific patches. GitLab CI/CD pipeline design including Kubernetes manifest validation gates, custom resource health checks, and environment promotion from lab through staging to production. |
Has operated a GitOps-exclusive delivery model in production where ArgoCD or Flux managed all cluster state and direct kubectl commands to production were prohibited. Designed a Kustomize overlay structure for a multi-cluster platform workload — not a single-cluster application. Built a GitLab CI pipeline with schema validation and promotion gates blocking invalid configurations before staging. Evidence of GitOps discipline applied to operator-managed custom resources, not only standard Deployments. |
| Networking | Intermediate | Kubernetes networking spanning Ingress controllers, LoadBalancer service patterns, CoreDNS configuration, and NetworkPolicy for multi-tenant namespace isolation. MetalLB deployment on bare-metal OpenShift in BGP mode: IPAddressPool and BGPAdvertisement CRD configuration, eBGP peering with upstream ToR and spine switches, AS number design, prefix filtering, route-map policy, and BFD fast-failover. OVN-Kubernetes SDN design including EgressIP, EgressFirewall, and cross-cluster traffic segmentation. |
Deployed MetalLB in BGP mode on bare-metal Kubernetes or OpenShift and configured eBGP peering with physical switching infrastructure — not Layer 2 mode only. Designed NetworkPolicy rules for cross-namespace and cross-cluster traffic paths. Has operated in an IPv6 or dual-stack network environment. Evidence of DNS design for platform services: split-horizon resolution, external DNS automation, and service record lifecycle management. |
| Telecommunications | Intermediate | Familiarity with the FCAPS operational framework — Fault, Configuration, Accounting, Performance, and Security — and how it shapes metric taxonomy, alerting rule design, and NOC-facing dashboard structure. Understanding of heterogeneous vendor telemetry integration: Prometheus exporter compatibility, OpenMetrics format validation, and label standardisation across multi-vendor network equipment from vendors such as Nokia and Ericsson. |
Has designed or operated an observability platform within a telecommunications, carrier-grade, or large-scale network infrastructure environment — 5G core, RAN, transport, or carrier edge. Experience normalising telemetry from heterogeneous vendor sources at the ingestion layer. Has integrated platform alerts into an enterprise NOC workflow including deduplication, suppression, and ticketing system integration. Prior engagement with a Tier-1 carrier, MVNO, or national-scale network operator strongly preferred. |
| Programme Delivery | Intermediate | Fluent in agile delivery practices within a structured programme: sprint planning, backlog refinement, and epic and story decomposition from architecture into engineering tasks, maintaining architecture decisions ahead of sprint execution. Able to work effectively across Product Management, Project Management, and Scrum Masters — translating technical design choices into programme risk, timeline, and trade-off language that non-technical stakeholders can act on. |
Has operated as an architect embedded in a governed programme with sprint ceremonies, formal review gates, and stakeholder reporting obligations — not a solo or small-team context. Authored Jira epics and stories that engineering teams executed without requiring repeated design clarification. Has presented a technical design decision to a Product Manager or Project Manager and adapted framing based on their concerns. Evidence of architecture governance discipline under delivery pressure: ADRs completed, gates honoured, exceptions documented. |
| Executive & Strategic Leadership | Expert | Developed and delivered enterprise IT strategy at Director level or above — translating business priorities into technology roadmaps, operating models, and investment cases for executive leadership. Experience leading large-scale digital transformation programmes spanning cloud adoption, platform modernisation, and operating model change. Partnered with C-Suite and SVP-level stakeholders to drive technology investment decisions, bridging business strategy and technology delivery across organisational boundaries. |
Has held a Director or equivalent role with enterprise-wide technology accountability — budget ownership above $50M, vendor governance, and SLA management across a global infrastructure estate. Built or led a Cloud Centre of Excellence delivering hybrid cloud governance and architectural standards. Led or governed a major managed services engagement (>$100M) including commercial model and operational improvement. Links technology investment to measurable business outcomes at programme scale. |
Benefits: Danta offers a compensation package to all W2 employees that are competitive in the industry. It consists of competitive pay, the option to elect healthcare insurance (Dental, Medical, Vision), Major holidays and Paid sick leave as per state law.
The rate/ Salary range is dependent on numerous factors including Qualification, Experience and Location.
Similar jobs
- SI
Controls Technical Training Specialist
NewStellent IT LLC
Grand Rapids, MI🇺🇸On-site19 hours agoConcrete - UG
Mammography Technologist I/II - Kelsey-Seybold Clinic: Springwoods Village
NewUnitedHealth Group
Spring, TX🇺🇸$29 - $52/hrHybrid19 hours ago401kComplianceData Entry+1 - JC
MS Dynamics Business Central
NewJC CORPORATIONS
Arizona City, AZ🇺🇸On-site19 hours agoERPProcurement - PA
Domain Operator (Life Sciences)- Contract to Hire- Remote (with travel to client sites)
NewPalni Inc
United States🇺🇸Remote19 hours agoStakeholder Management - YT
AIX / UNIX Lead Admin
NewYashnee Tech Solutions Corporation
Albany, NY🇺🇸On-site19 hours agoShellC#.NET - RS
UAT/QE Lead
NewRuri Software Technologies LLC
Newark, NJ🇺🇸Hybrid19 hours agoCapacity PlanningComplianceInternal Audit+3