Quick Overview
Job Description
Position Title: Kubernetes (K8s) / AI Platform Architect – On-Premises
Location: Milpitas, CA
Experience
12+ Years in IT Infrastructure, Platform Engineering, Cloud-Native Technologies, and Enterprise Architecture
5+ Years in Kubernetes Platform Architecture and Operations
3+ Years in AI/ML Platform Design and Enablement
Job Summary
We are seeking a highly experienced Kubernetes / AI Platform Architect to design, implement, and govern enterprise-grade on-premises AI and Kubernetes platforms supporting Generative AI, Machine Learning, Data Science, and cloud-native workloads.
The ideal candidate will lead the architecture of scalable, secure, GPU-enabled Kubernetes environments across private data centers. This role will be responsible for defining platform standards, AI infrastructure strategies, MLOps capabilities, container orchestration frameworks, and operational excellence practices to support business-critical AI and application workloads.
Key Responsibilities
AI Platform Architecture
- Define enterprise AI platform architecture supporting GenAI, LLMs, Machine Learning, and Data Science workloads.
- Design GPU-accelerated infrastructure for AI model training, fine-tuning, and inferencing.
- Establish scalable AI platform blueprints and reference architectures.
- Evaluate and integrate AI frameworks, model serving solutions, vector databases, and inference platforms.
- Enable self-service AI environments for data scientists and ML engineers.
Kubernetes Platform Architecture
- Design and implement enterprise-grade Kubernetes platforms on-premises.
- Define multi-cluster architecture, cluster lifecycle management, and multi-tenancy strategies.
- Build highly available, resilient, and scalable container platforms.
- Establish Kubernetes standards, governance, security, and operational models.
- Design platform services supporting modern cloud-native applications.
Platform Engineering & Automation
- Architect Infrastructure as Code (IaC) and platform automation frameworks.
- Implement GitOps practices utilizing ArgoCD, FluxCD, or equivalent technologies.
- Automate cluster deployment, patching, upgrades, and configuration management.
- Establish CI/CD pipelines supporting application and AI workload deployment.
MLOps & AI Operations
- Design MLOps workflows supporting model development, deployment, monitoring, and retraining.
- Implement ML lifecycle management and model governance practices.
- Enable automated model serving and inferencing pipelines.
- Define observability frameworks for AI workloads.
- Support responsible AI governance and deployment standards.
Security & Compliance
- Design secure Kubernetes environments using Zero Trust principles.
- Implement container image security, vulnerability management, and runtime protection.
- Architect identity and access management frameworks.
- Define network segmentation and micro-segmentation controls.
- Ensure compliance with enterprise security and regulatory requirements.
Infrastructure & Data Center Integration
- Integrate Kubernetes with enterprise storage and networking platforms.
- Design GPU resource allocation and optimization frameworks.
- Enable seamless integration with virtualization platforms such as VMware/OpenShift ecosystems.
- Define disaster recovery and business continuity architectures.
Leadership & Stakeholder Management
- Lead platform strategy discussions with customer leadership and enterprise architects.
- Mentor Platform Engineers, DevOps Engineers, SREs, and AI Engineers.
- Drive technical governance and architecture review boards.
- Collaborate with business, infrastructure, security, and application teams.
Required Technical Skills
Kubernetes Ecosystem
- Kubernetes (K8s)
- OpenShift
- Rancher
- Tanzu
- KubeVirt
- Helm
- Operators
- Container Runtime Technologies
Container Technologies
- Docker
- Podman
- Container Registry Solutions
- Container Security Platforms
AI & Machine Learning Platforms
- Kubeflow
- MLflow
- Ray
- JupyterHub
- NVIDIA AI Enterprise
- TensorFlow
- PyTorch
- Hugging Face
- Model Serving Platforms
GPU & AI Infrastructure
- NVIDIA GPUs
- CUDA
- MIG (Multi-Instance GPU)
- GPU Operator
- GPU Scheduling
- AI Cluster Design
DevOps & GitOps
- GitHub
- GitLab
- Jenkins
- ArgoCD
- FluxCD
- Tekton
- Terraform
- Ansible
Observability & SRE
- Prometheus
- Grafana
- Loki
- ELK Stack
- OpenTelemetry
- Jaeger
Networking
- Cilium
- Calico
- Service Mesh (Istio/Linkerd)
- Load Balancers
- Storage Networking
Storage Platforms
- Ceph
- Portworx
- NetApp
- Dell PowerScale
- Persistent Volume Management
Security
- HashiCorp Vault
- RBAC
- OPA/Gatekeeper
- Kyverno
- Prisma Cloud
- Aqua Security
- Falco
Required Qualifications
- Bachelor’s or master’s degree in computer science, Engineering, Information Technology, or related discipline.
- 12+ years of overall infrastructure, cloud, platform engineering, or architecture experience.
- 5+ years designing and implementing Kubernetes platforms.
- Hands-on experience with enterprise container platforms in production environments.
- Experience designing GPU-enabled infrastructure for AI/ML workloads.
- Strong understanding of platform engineering, DevOps, GitOps, and SRE practices.
- Experience supporting mission-critical enterprise environments.
Preferred Certifications
- Certified Kubernetes Administrator (CKA)
- Certified Kubernetes Application Developer (CKAD)
- Certified Kubernetes Security Specialist (CKS)
- Red Hat OpenShift Administration Certification
- NVIDIA AI Infrastructure Certification
- VMware Tanzu Certification
- HashiCorp Terraform Certification
Preferred Experience
- Enterprise GenAI and LLM deployment initiatives.
- AI Factory or Internal AI Platform implementations.
- GPU cluster architecture and optimization.
- Large-scale Kubernetes environments (5000+ containers).
- Hybrid cloud and multi-cloud architecture.
- Platform Engineering Center of Excellence initiatives.
- FinOps and infrastructure cost optimization.
- Telecom, Financial Services, Healthcare, Manufacturing, or Hyperscale environments.
Key Competencies
Technical Leadership
- Enterprise Architecture
- Kubernetes Strategy
- AI Platform Design
- Platform Engineering
Operational Excellence
- Site Reliability Engineering (SRE)
- Automation First Mindset
- Incident Reduction
- Reliability Engineering
Business & Communication
- Customer Engagement
- Executive Communication
- Stakeholder Management
- Strategic Roadmap Development
Governance & Risk
- Security Architecture
- Compliance Management
- Platform Governance
- Risk Mitigation
Success Measures
- Successful deployment of enterprise-grade Kubernetes platforms.
- Reliable and scalable AI/ML platform adoption across business units.
- Reduction in infrastructure provisioning time through automation.
- Increased GPU utilization and infrastructure efficiency.
- Improved developer and data scientist productivity.
- High availability and resilience of critical workloads.
- Improved platform security posture and compliance adherence.
- Successful onboarding of AI and cloud-native workloads.
“Tech Mahindra is an Equal Employment Opportunity employer. We promote and support a diverse workforce at all levels of the company. All qualified applicants will receive consideration for employment without regard to race, religion, color, sex, age, national origin, or disability. All applicants will be evaluated solely on the basis of their ability, competence, and performance of the essential functions of their positions with or without reasonable accommodations. Reasonable accommodations also are available in the hiring process for applicants with disabilities. Candidates can request a reasonable accommodation by contacting the company ADA Coordinator at .”
Similar jobs
- TG
SAP TSW(Trader's and Scheduler's Workbench)
NewTalent Groups
Houston, TX🇺🇸On-siteYesterdayTechnology - PG
Release and Configuration Manager, Hybrid- 70682
NewPRIMUS Global Services Inc.
Michigan City, IN🇺🇸HybridYesterdayPayrollTechnology - SP
Immediate Opening for Desktop Support in Chicago (Downtown) IL- Onsite || Contract
NewSR Partners LLC
Chicago, IL🇺🇸On-siteYesterdayAWSActive DirectoryAzure+7Technology - TS
Salesforce DevOps / Release Engineer
NewTechSpace Solutions Inc.
Plano, TX🇺🇸On-siteYesterdayGitJiraRESTTechnology - EG
Java Backend Developer
NewEnexus Global
Pittsburgh, PA🇺🇸HybridYesterdayDockerMicroservicesMongoDB+13Technology - SI
PA - Security Information and Event Management (SIEM) Engineer - 813707
NewSR International Inc.
Harrisburg, PA🇺🇸HybridYesterdaySplunkTechnology - S&
PeopleSoft Developer
NewSyntrux & Co
United States🇺🇸HybridYesterdaySOAPSQLRESTTechnology - CB
Cybersecurity Solutions Architect – AI (Insurance Industry) - AS
NewCentral Business Solutions
United States🇺🇸HybridYesterdayAWSEncryptionSOC 2+8Technology - HP
Technical Program Manager_ SAP Signavio
NewHR Pundits
United States🇺🇸RemoteYesterdayAgileERPLean Six Sigma+4Technology - KT
Integration Architect
NewKforce Technology Staffing
Denver, CO🇺🇸HybridYesterdayETLAgileAuditing+4Technology - MP
Solutions Architect – Java
NewMpower Plus Rezolve AI Group LTD
Sunnyvale, CA🇺🇸HybridYesterdayMongoDBOracleAgile+2Technology - KT
Data Engineer
NewKforce Technology Staffing
Denver, CO🇺🇸HybridYesterdaySQLETLAgile+3Technology