Quick Overview
Job Description
Mirantis, an IREN company, is the Kubernetes-native AI infrastructure company, enabling organizations to build and operate scalable, secure, and sovereign infrastructure for modern AI, machine learning, and data-intensive applications. By combining open source innovation with deep expertise in Kubernetes orchestration, Mirantis empowers platform engineering teams to deliver composable, production-ready developer platforms across any environment—on-premises, in the cloud, at the edge, or in sovereign data centers. As enterprises navigate the growing complexity of AI-driven workloads, Mirantis delivers the automation, GPU orchestration, and policy-driven control needed to manage infrastructure with confidence and agility. Committed to open standards and freedom from lock-in, Mirantis ensures that customers retain full control of their infrastructure strategy. https://www.mirantis.com/
We are looking for a Senior Data Platform Engineer to run the PostgreSQL and Apache Kafka platform behind k0rdent-ai — our multi-tenant control plane for enterprise GPU infrastructure. Every cluster provisioned, every GPU-hour consumed, and every tenant action is recorded on this platform, so it must be reliable, secure, and recoverable.
This is a hands-on DevOps / SRE role for stateful systems. You will deploy and operate PostgreSQL (CloudNativePG) and Kafka (Strimzi) on Kubernetes using operators, Helm, and GitOps — across cloud and bare-metal clusters, in a global control plane and multiple regions. You will own automation, upgrades, observability, backups, and disaster recovery, and give product teams self-service access to databases and topics through code.
Main Responsibilities
- Run PostgreSQL and Kafka on Kubernetes: Deploy, upgrade, scale, and operate CloudNativePG and Strimzi clusters with Helm charts and operator custom resources, across cloud and bare-metal Kubernetes.
- GitOps & Infrastructure as Code: Manage everything declaratively with Argo CD (or Flux), Helm, and Terraform; build CI/CD pipelines for database and Kafka changes, schema migrations, and operator upgrades.
- High Availability & Disaster Recovery: Design and operate high availability within each region and cross-region disaster recovery — PostgreSQL replica clusters, backups and point-in-time recovery, Kafka MirrorMaker 2 — with defined RPO/RTO targets and regular failover drills.
- Self-Service Platform: Give product teams self-service provisioning of databases, users, topics, and ACLs through code, with secure defaults and guardrails.
- Observability & Operations: Build monitoring and alerting (Prometheus, Grafana, OpenTelemetry) for replication lag, backups, consumer lag, and capacity; participate in on-call and lead incident reviews.
- Security & Multi-Tenancy: Implement TLS/mTLS, secrets management (Vault/OpenBao), network policies, Kafka ACLs, and per-tenant isolation; integrate database and Kafka access with our identity platform (Keycloak / OIDC).
- Data Movement & CDC: Operate the outbox and change-data-capture pipelines (Debezium, Kafka Connect) that keep databases and event streams consistent between the global control plane and regions.
- Automation & Tooling: Write automation and small services in Go or Python to remove manual work, and partner with application teams on schema standards and performance troubleshooting.
Must Have
- Experience: 8+ years in DevOps, SRE, platform, or infrastructure engineering, including 3+ years running stateful systems (databases or message streaming) in production.
- Kubernetes: Strong hands-on Kubernetes operations — StatefulSets, persistent volumes and storage classes, pod disruption budgets, scheduling and affinity, network policies, and cluster upgrades — on cloud and/or bare-metal clusters.
- Operators & Helm: Production experience running PostgreSQL and/or Kafka through Kubernetes operators (CloudNativePG, Zalando, or Crunchy PGO; Strimzi, Confluent for Kubernetes, or similar), and authoring and maintaining Helm charts. Depth with one operator matters more than the specific product.
- GitOps & IaC: Day-to-day use of Argo CD or Flux, Terraform, and CI/CD pipelines (GitHub Actions, GitLab CI, or similar) in a declarative, review-driven workflow.
- PostgreSQL Operations: Practical PostgreSQL administration — replication and failover, backup and point-in-time recovery, connection pooling (PgBouncer), version upgrades, and performance troubleshooting.
- Kafka Operations: Running Kafka in production — brokers, topics and partitions, replication, consumer groups, monitoring, and capacity planning.
- Reliability & DR: Designing and testing high availability and disaster recovery for stateful systems, with defined RPO/RTO.
- Automation: Strong scripting and programming in Go (preferred) or Python, plus Bash.
Nice to Have
- Cross-region replication (CloudNativePG replica clusters, Kafka MirrorMaker 2) and global/regional multi-site architectures.
- Debezium, Kafka Connect, and transactional outbox / CDC patterns.
- Identity integration for data platforms (Keycloak, OIDC, LDAP/Active Directory).
- Observability stacks (Prometheus, Grafana, OpenTelemetry, VictoriaMetrics).
- Secrets and security tooling (Vault/OpenBao, cert-manager, mTLS).
- Bare-metal infrastructure, Ceph or object storage, or GPU/AI infrastructure.
- Contributions to CloudNativePG, Strimzi, or related open-source projects.
- Experience in SOC 2 or ISO 27001 environments.
Education and Experience
Bachelor's degree in Computer Science or a related field, or equivalent practical experience.
What does Mirantis offer you?
- Work with an established Silicon Valley leader in the cloud infrastructure industry;
- Work with exceptionally passionate, talented and engaging colleagues, helping Fortune 500 and Global 2000 customers implement next-generation cloud technologies;
- Be a part of cutting-edge, open-source innovation;
- Thrive in the high-energy environment of a young company where openness, collaboration, risk-taking, and continuous growth are valued;
- Professional development and training;
- Attend conferences and working groups;
- Company outings, happy hours, hackathons, and tech talks;
- Receive a competitive compensation package with a strong benefits plan.
We are a Leader for Container Management in G2 (#2 after AWS)!
Similar jobs
- MS
Senior DevSecOps Engineer
NewAuto ApplyMomentus Space LLC
San Jose🇺🇸Hybrid4 hours agoShellAWSRobotics+11Engineering - LI
DevOps Engineer
NewAuto ApplyLIGHTFEATHER IO LLC
United States🇺🇸Hybrid8 hours agoDockerRubyAWS+4Technology - LI
Staff Software Engineer (DevOps / Platform Engineering)
NewAuto ApplyLiberate
Boston / San Francisco🇺🇸Hybrid9 hours agoAWSComplianceGitHub Actions+8Technology - FO
DevSecOps Engineer - Leesburg
NewAuto ApplyFortreum
Leesburg🇺🇸Hybrid9 hours agoGCPAWSSplunk+13Engineering - FO
DevSecOps Engineer - Reston
NewAuto ApplyFortreum
Reston🇺🇸Hybrid9 hours agoGCPAWSSplunk+13Engineering - CL
Senior Platform Engineer
NewAuto ApplyClera
San Francisco🇺🇸On-site4 hours agoGCPAWSKubernetes+3Technology - EN
Senior DevOps Engineer
NewAuto ApplyEncoura
Remote🇺🇸Remote11 hours agoDockerMongoDBSQL+14Technology - TA
Staff Infrastructure Engineer
NewAuto ApplyTabs
New York City🇺🇸On-site8 hours agoDockerAWSCompliance+8Technology - RE
Platform Engineer
NewAuto ApplyResend
Americas🇺🇸Remote10 hours agoNode.jsAWSLoad Balancing+6Technology - 2K
Senior Site Reliability Engineer
NewAuto Apply2K
Austin🇺🇸Hybrid9 hours agoMySQLPackerAWS+18Technology - OP
Site Reliability Engineer, Provider Operations
NewAuto ApplyOpenrouter
Remote (US)🇺🇸Remote7 hours agoGCPAPI GatewayLoad Balancing+7Technology - TH
Senior DevOps Engineer
NewAuto ApplyTheIncLab
Colorado Springs, Colorado🇺🇸HybridYesterdayAWSAnsibleBash+12Technology