Quick Overview
Seniority
Mid Senior
Work mode
Remote
Location
United States
Posted
18 hours ago
SQLScalaETLEncryptionMLOpsMLflowMachine LearningSSOAzureComplianceDatabricksGitHub ActionsKafkaPythonRESTRegulatory ReportingStakeholder ManagementTerraformUnityVaultpytest
Job Description
Role: Databricks Solution Expert (existing or past VA experience preferred)
Location: 100% Remote
Job Description:
Responsibilities:
The Databricks Solutions Expert will define platform patterns, build reference implementations, set standards (cluster policies, Unity Catalog governance, CI/CD), and drive enablement so product, analytics, and data science teams can deliver reliable, compliant, and performant insights at scale.
Platform Architecture & Design
Blueprint the Azure Databricks landing zone: workspace topology (prod/non-prod), network architecture (VNet injection, Private Link, NAT), secure connectivity to ADLS Gen2, Azure SQL/MI, Event Hubs, and other data sources.
Governance with Unity Catalog: multi-region metastores, catalog/schema/table design, data classification tiers, row/column-level security, data lineage, and cross-domain data sharing patterns (Delta Sharing).
Lakehouse foundations: Delta Lake storage design (bronze/silver/gold), medallion data flow standards, partitioning, Z‑ordering, OPTIMIZE/VACUUM policies, and performance best practices (Photon, AQE).
Scalability for “thousands of teams”: multi-workspace strategy, tenancy model (shared vs. dedicated), workspace baselines, cluster policy tiers, and guardrails to prevent noisy-neighbor and cost runaways.
Reliability: HA/DR strategy, regional deployments, backup/restore, versioning, and repeatable environment provisioning via Terraform.
Security, Compliance & Access Control
Identity & access: Entra ID (Azure AD) SSO, SCIM user/group provisioning, service principals/managed identities, attribute-based/role-based access controls mapped to Unity Catalog.
Secrets & credentials: Key Vault–backed secret scopes, credential passthrough patterns, token hygiene (PAT governance).
Data protection: encryption at rest/in transit, private endpoints, data exfil and egress controls, policy-as-code, and audit logging to Log Analytics or secure storage.
Compliance: enforce enterprise policies (PII/PHI handling), data retention, legal hold, and regulatory reporting with auditable lineage.
Engineering & Enablement
Pipelines: build DLT (Delta Live Tables) pipelines and jobs for batch and streaming (Structured Streaming) with CDC (e.g., via Auto Loader) from enterprise sources.
Performance engineering: optimize notebooks/SQL/ETL (Photon, caching, skew mitigation), tune cluster sizing, and set standards for reliable, fast jobs.
MLOps: integrate MLflow for experiment tracking, model registry, feature store (Unity Catalog), and serving patterns where appropriate.
Observability: end-to-end monitoring (jobs, clusters, UC audits), dashboarding, alerting, and SLOs; integrate with Azure Monitor/Log Analytics.
Enablement: build reusable reference accelerators (templates, example notebooks, data products), run playbooks, and conduct office hours to uplevel 1000s of teams.
DevOps, Automation & Cost Management
Infrastructure-as-Code: provision workspaces, catalogs, cluster policies, and permissions via Terraform (Databricks provider), with pipelines in Azure DevOps/GitHub.
CI/CD: notebook/package deployment, testing harnesses (dbx/pytest), environment promotion, and artifact versioning.
FinOps: cost modeling, budgets/alerts, instance pools, auto-termination, serverless SQL, tagging/chargeback, and usage analytics for executive reporting.
Stakeholder Management & Governance
Partner with Security, Networking, Compliance, and FinOps to codify enterprise standards.
Establish a Lakehouse Platform Council to ratify patterns and review exceptions.
Create adoption metrics, business case narratives, TCO models, and executive updates.
Required Skills and Experience:
Bachelor’s degree in Information Technology or a related field. (or commensurate experience)
12+ years of experience in data engineering/analytics; 5+ years building on cloud data platforms (Azure preferred).
3+ years hands-on Azure Databricks (platform + pipelines) and Delta Lake.
Proven experience setting up Unity Catalog with granular governance (RLS/CLS).
Deep knowledge of Azure networking (VNets, Private Link, NSGs), identity (Entra ID), Key Vault, ADLS Gen2, Event Hubs, Azure SQL/MI, and Data Factory.
Strong Spark expertise (PySpark/SQL), Structured Streaming, performance tuning, partitioning and storage optimization.
Practical Terraform experience for Databricks/Azure resources; CI/CD with Azure DevOps or GitHub Actions.
Security-first mindset; track record implementing audit logging, policy-as-code, and compliance controls.
Excellent communication skills; ability to standardize, teach, and influence at enterprise scale.
Ability to obtain and maintain a Suitability/Public Trust clearance
Preferred Skills and Experience
Experience operating platforms for >500 concurrent users and 1000s of analytics teams.
Knowledge of Photon, DLT, Workflows, Lakehouse ML (MLflow, feature store), Delta Sharing.
Exposure to FinOps and chargeback models for data platforms.
Experience with Synapse/Fabric interoperability, and data virtualization patterns.
Background in data modeling (medallion, dimensional, domain‑driven design) and data quality (expectations, SLAs).
Certifications
Databricks: Databricks Certified Data Engineer Professional, Lakehouse Fundamentals, Machine Learning Associate/Professional.
Microsoft: Azure Solutions Architect Expert (AZ‑305), Azure Data Engineer (DP‑203), Azure Security Engineer (AZ‑500).
HashiCorp: Terraform Associate.
Technical Stack & Tools
Core: Azure Databricks (Unity Catalog, Delta Lake, DLT, Workflows, Photon), ADLS Gen2, Azure Key Vault, Entra ID, Private Link.
Data Integration: Auto Loader, ADF/Synapse pipelines, Event Hubs/Kafka, JDBC/ODBC.
DevOps/Infra: Terraform (Databricks & Azure providers), Azure DevOps/GitHub Actions, dbx.
Observability: Databricks audit logs, Azure Monitor, Log Analytics, custom usage analytics.
Languages: Python (PySpark), SQL, Scala (optional).
ML: MLflow, Feature Store (Unity Catalog), model serving patterns.
Similar jobs
- GL

Step Into Leadership. Build Your Business. Lead Your Team.
GIA Legacy Planning
Kingsville🇺🇸Remote4 days agoCRM - GL

Make an Impact. Build a Career with Purpose.
GIA Legacy Planning
Carrollton🇺🇸Remote4 days agoCRM - GL

Step Into Leadership. Build Your Business. Lead Your Team.
GIA Legacy Planning
Denham Springs🇺🇸Remote1 week agoCRM - GL

💥 Make an Impact. Build a Career with Purpose.
GIA Legacy Planning
Irving🇺🇸Remote1 week agoCRM - GL

💥 Make an Impact. Build a Career with Purpose.
GIA Legacy Planning
San Antonio🇺🇸Remote2 weeks agoCRM - GL

🚀 Join Our Team of Final Expense Agents!
GIA Legacy Planning
Columbia🇺🇸Remote2 weeks agoCRM