Haystack
← Back to Jobs
Other
UI

Databricks Architect

Unikon ITUnited States🇺🇸United StatesPosted Sep 21, 2026

Quick Overview

Seniority
Mid Senior
Work mode
Hybrid
Location
United States
Posted
20 hours ago
DockerSQLScalaAWSETLMLOpsMLflowMachine LearningApacheApache SparkAuditingAzureComplianceDatabricksGenerative AIGitHub ActionsGoogle CloudHIPAAJavaPythonTerraformUnity

Job Description

Job Title: Databricks Architect
Location: Remote 


Role Summary
We are seeking an expert Principal Databricks Architect to lead the vision, design, and implementation of our next-generation enterprise Data Intelligence Platform. In this role, you will be the senior technical authority and trusted advisor responsible for building highly secure, scalable, and cost-effective data infrastructure across multi-cloud environments (Azure, AWS, and Google Cloud Platform).
You will design robust Lakehouse architectures, implement enterprise-grade data governance, automate end-to-end ETL/ELT pipelines, and architect cutting-edge GenAI and machine learning infrastructure. This position requires a deep technical understanding of distributed systems, advanced cybersecurity controls, HIPAA compliance, and state-of-the-art MLOps/DataOps workflows to align technical strategy with business outcomes.

Key Responsibilities
1. Data Platform Architecture & Engineering
  • Lakehouse & Medallion Design: Architect and scale an enterprise-wide Databricks Lakehouse leveraging the Medallion Architecture (Bronze, Silver, Gold) to serve operational, analytics, and AI workloads.
  • Pipeline Automation: Design and optimize real-time and batch data pipelines using Delta Live Tables (DLT), Apache Spark, PySpark, and Spark SQL.
  • Storage Optimization: Enforce Delta Lake best practices, maximizing performance via transaction logs, ACID compliance, data compaction (OPTIMIZE), multi-dimensional clustering (Z-ORDER), and Auto Loader.
  • Application Ecosystem: Oversee the development and integration of Databricks Apps to deliver data-driven tools directly within the platform ecosystem.
2. Enterprise Governance, Security & Compliance
  • Unity Catalog Enforcement: Standardize data governance across the enterprise by implementing Unity Catalog for centralized access management, data lineage tracking, and auditing.
  • Granular Security Controls: Design and automate advanced security rules, including Row-Level and Column-Level Security as well as dynamic data masking to safeguard sensitive assets.
  • Compliance & RBAC: Build Automated Role-Based Access Control (RBAC) matrices integrated with cloud identity management. Ensure the entire platform architecture strictly complies with HIPAA compliance regulations for protective health information.
  • Observability & Quality Gates: Implement automated Data Quality Gates to stop bad data from propagating, while monitoring operational health with advanced observability and auditing dashboards.
3. AI, Generative AI & MLOps
  • RAG & GenAI Integration: Architect semantic search and generative AI pipelines using Databricks Vector Search, Azure OpenAI, Retrieval-Augmented Generation (RAG) frameworks, and customized Vector Embeddings.
  • ML Lifecycle Management: Standardize machine learning workflows across the organization by deploying MLflow for model tracking, versioning, and environment reproducibility.
  • MLOps Pipelines: Collaborate with data scientists to establish production-grade MLOps practices, automating model deployment, monitoring, and retraining pipelines.
4. Infrastructure, DevOps & Cloud Integration
  • Multi-Cloud Ecosystems: Integrate Databricks seamlessly with cloud provider stacks, including Azure (Azure Databricks, ADLS Gen2, Synapse, Azure Data Factory), AWS, and Google Cloud Platform.
  • Infrastructure as Code (IaC): Automate the orchestration, deployment, and configuration of Databricks workspaces and clusters using Terraform or other IaC tools.
  • CI/CD Automation: Establish strict CI/CD and DataOps release engineering pipelines for notebooks, workflows, and infrastructure using Azure DevOps or GitHub Actions.
  • Containerization: Utilize Docker to containerize application environments, testing frameworks, and custom execution environments.
5. FinOps & Platform Optimization
  • Compute Tuning: Maximize system performance and reduce latency through targeted cluster optimization, custom cluster sizing, auto-scaling thresholds, Photon acceleration, and proactive instance selection.
  • FinOps Governance: Define organizational cluster policies, budget alerts, and cost allocation models to eliminate wasted cloud expenditure.

Required Skills & Qualifications
  • Experience: 12+ years in Data Engineering/Architecture, with 3+ years of dedicated experience architecting enterprise-scale Databricks environments.
  • Core Coding: Expert-level mastery of Python, PySpark, and complex SQL optimization (proficiency in Scala or Java is a plus).
  • Data Security Expertise: Proven track record implementing Unity Catalog, Row/Column-level security, and building environments governed by HIPAA or similar compliance frameworks.
  • Cloud Infrastructure: Strong background managing native storage and orchestrations across Azure (ADLS, ADF), AWS, or Google Cloud Platform.
  • Modern Tools: Hands-on experience with Docker, Terraform, Databricks Workflows, and CI/CD pipeline automation.
  • AI Frameworks: Practical architectural understanding of LLMs, Vector Databases, and MLflow.
  • Certifications (Preferred): Databricks Certified Data Engineer or Platform Architect credentials

Similar jobs