Haystack
← Back to Jobs
Other
KA

Lead Databricks Architect

KaltechsoftIndianapolis, IN🇺🇸United StatesPosted Oct 5, 2026

Quick Overview

Seniority
Mid Senior
Work mode
Hybrid
Location
Indianapolis, IN, United States
Posted
Yesterday
DynamoDBSQLAWSETLCloudFormationComplianceData PipelineDatabricksGitHub ActionsPostgreSQLPythonTerraformUnityVault

Job Description

Job Title: Lead Databricks Architect
Location: Hybrid, Indianapolis
Industry: Pharmaceutical
Job Overview:
Theoris is assisting our client in their search for a Lead Databricks Architect to lead the architecture, governance, and modernization of an enterprise Databricks platform running on AWS.
This is a true architecture and platform leadership role, not a general Databricks development or data engineering position. The ideal candidate has already led multiple enterprise-scale Databricks implementations on AWS and has direct experience migrating AWS Glue, Aurora PostgreSQL, and existing AWS data lake environments into a Databricks Lakehouse architecture.
The architect will establish platform standards and reusable architectural patterns while guiding migration strategy, governance, security, and engineering practices across multiple Databricks workspaces in a large, regulated enterprise environment.
Responsibilities:

  • Design and build metadata-driven, config-based raw-to-refine ingestion pipelines for large-scale Veeva Vault object sets (e.g., 200+ objects) using Databricks on AWS with a Unity Catalog medallion (raw/refine) architecture.
  • Develop and maintain Databricks Lakeflow Declarative Pipelines (formerly DLT), including CDC decorator patterns for inserting/update/delete tracking and materialized views for relationship expansion.
  • Integrate with the Veeva Direct Data API (DDAPI) and/or Databricks Lakeflow Veeva Vault Connector, including managing incremental file timing, schema evolution, and field-rename edge cases.
  • Write and optimize Python-based code generators and automation utilities that translate metadata exports into production pipeline definitions (e.g., SQL/MV templates).
  • Author advanced SQL for Databricks/Delta and Iceberg-compatible tables, resolving platform-specific issues (e.g., LATERAL VIEW/JOIN chaining, correlated subqueries, deletion vectors, REORG PURGE).
  • Define and provision cloud and platform infrastructure using Terraform, including Databricks workspaces, Unity Catalog objects, service principals, and IAM/permissions.
  • Build and maintain AWS infrastructure using CloudFormation (CFN), integrating with services such as S3, Glue, Lambda, Step Functions, Event Bridge, and DynamoDB where applicable.
  • Design manifest-driven CI/CD approval and deployment workflows (e.g., GitHub Actions) leveraging the Databricks Statement Execution API for controlled production releases.
  • Ensure pipelines meet the compliance, traceability, and data-quality standards required in regulated pharma domains, understanding Veeva Vault as the system of record for transactional data.
  • Partner with Safety/Regulatory IT teams, data architects, and business stakeholders to translate Veeva data domain requirements into scalable engineering solutions.
  • Troubleshoot and resolve complex data pipeline, schema evolution, and platform compatibility issues; drive root-cause fixes and long-term architectural improvements.
  • Document architecture decisions, naming conventions, and operational runbooks; mentor other engineers on Veeva/Databricks best practices.
  • Serve as the lead architect for an enterprise Databricks platform hosted on AWS.
  • Define Databricks platform architecture, technical standards, governance models, security controls, and engineering guardrails.
  • Establish standards across multiple Databricks workspaces, including workspace design, Unity Catalog, data sharing, access controls, and platform governance.
  • Lead migrations from AWS Glue, Aurora PostgreSQL, and existing AWS data lake environments to Databricks.
  • Modernize existing ETL frameworks, data pipelines, and operational processes as part of the migration strategy.
  • Develop reusable platform accelerators, reference architectures, deployment patterns, and engineering standards.
  • Design and implement enterprise data governance utilizing Unity Catalog and appropriate RBAC/ABAC models.
  • Partner with security, infrastructure, data engineering, architecture, and business teams to ensure platform solutions meet enterprise security, compliance, privacy, and audit requirements
  • Establish CI/CD and Infrastructure-as-Code practices for Databricks environments.
  • Provide hands-on technical leadership and architectural guidance throughout implementation and migration efforts.

 
 Requirements

  • Extensive Databricks architecture experience within AWS environments.
  • Proven experience leading multiple enterprise Databricks implementations.
  • Experience architecting Databricks environments containing multiple workspaces.
  • Direct experience leading AWS Glue-to-Databricks migrations.
  • Experience migrating Aurora PostgreSQL and/or PostgreSQL-based workloads into Databricks.
  • Experience modernizing AWS data lake environments into Databricks Lakehouse architectures.
  • Deep expertise with Unity Catalog, Delta Lake, Databricks Workflows, SQL Warehouses, and Spark/PySpark.
  • Experience with declarative data pipeline frameworks.
  • Strong AWS knowledge including S3, IAM, networking, and security controls.
  • Experience implementing CI/CD and Infrastructure as Code within enterprise data platforms.
  • Demonstrated experience defining platform standards, reference architectures, governance frameworks, and migration strategies.

 Preferred

  • Experience within pharmaceutical, biotech, healthcare, financial services, or another highly regulated enterprise environment.
  • Experience supporting audit, privacy, security, and compliance requirements.
  • Knowledge of GxP environments and validated systems.
  • Experience implementing enterprise RBAC/ABAC and data governance controls
  •  

Similar jobs