Haystack
← Back to Jobs
Other
TT

Databricks SME

Tech3pillars TechnologiesUnited States🇺🇸United StatesPosted 3 Sept 2026

Why This Role Stands Out

This hybrid role offers a fantastic opportunity to deepen your expertise in Databricks administration and production support within a reputable tech company. You'll thrive here if you're a proactive problem-solver passionate about maintaining high-performing data platforms, and we encourage you to apply!

Quick Overview

Seniority
Mid Senior
Work mode
Hybrid
Location
United States
Posted
1 week ago
SQLAWSApacheApache SparkAzureCapacity PlanningDatabricksGitGoogle CloudPythonRoot Cause AnalysisTerraformUnity

Job Description

Job Title: Databricks SME – Operations & Administration
Work Arrangement: Remote

Job Summary

We are seeking an experienced Databricks Subject Matter Expert (SME) with strong hands-on experience in Databricks platform administration, production support, troubleshooting, configuration, and operational activities. The ideal candidate will be responsible for maintaining the stability, performance, availability, and security of the Databricks environment and supporting critical production workloads.

Key Responsibilities

  • Provide L2/L3 operational support and administration for enterprise Databricks environments.
  • Perform day-to-day Databricks platform administration, configuration, and troubleshooting.
  • Manage and troubleshoot Databricks Workspaces, Clusters, Jobs, Notebooks, Workflows, and SQL Warehouses.
  • Monitor production workloads and proactively identify and resolve platform and job-related issues.
  • Troubleshoot cluster failures, job failures, performance issues, connectivity issues, and resource-related problems.
  • Configure and manage cluster policies, compute resources, runtime versions, libraries, and access controls.
  • Manage Databricks Jobs/Workflows, schedules, dependencies, alerts, and failed executions.
  • Perform root cause analysis (RCA) for recurring production issues and implement permanent fixes.
  • Monitor platform utilization, cluster performance, job execution, and resource consumption.
  • Support capacity planning, performance tuning, and optimization of Databricks environments.
  • Administer users, groups, permissions, service principals, tokens, and workspace access.
  • Implement and troubleshoot RBAC and security configurations in Databricks.
  • Support integration with Azure/AWS cloud services, storage, networking, and enterprise data platforms.
  • Troubleshoot connectivity between Databricks and external systems such as ADLS/S3, databases, APIs, data warehouses, and messaging platforms.
  • Support Databricks Runtime upgrades, patches, configuration changes, and platform maintenance.
  • Coordinate with Cloud, Data Engineering, Security, Network, and Application teams for production incidents.
  • Participate in incident, problem, and change management processes.
  • Maintain operational documentation, runbooks, knowledge articles, and troubleshooting procedures.
  • Participate in on-call/after-hours production support as required.

Required Skills

  • Strong hands-on experience with Databricks administration and production operations.
  • Excellent troubleshooting skills across Databricks platform, clusters, jobs, workflows, and connectivity.
  • Strong knowledge of Databricks Workspaces, Compute/Clusters, Jobs, Workflows, Delta Lake, and Databricks SQL.
  • Experience with cluster configuration, cluster policies, autoscaling, libraries, and runtime management.
  • Experience with Databricks security, RBAC, users, groups, service principals, and access management.
  • Strong knowledge of Apache Spark and Spark troubleshooting.
  • Experience with Python and/or PySpark, SQL, and scripting.
  • Experience with at least one major cloud platform: Azure, AWS, or Google Cloud Platform.
  • Strong knowledge of cloud storage such as ADLS Gen2, Amazon S3, S.
  • Experience with monitoring, alerting, logging, and production incident management.
  • Strong understanding of performance tuning and capacity management.
  • Excellent communication and problem-solving skills.

Preferred / Nice to Have

  • Experience with Databricks on Azure (Azure Databricks).
  • Experience with Unity Catalog and data governance.
  • Experience with Terraform/IaC for Databricks configuration.
  • Knowledge of Azure DevOps/Git and CI/CD.
  • Experience with Databricks APIs and CLI.
  • Experience supporting large-scale enterprise production environments.
  • Knowledge of ITIL, incident management, change management, and RCA processes.

Ideal Candidate: A hands-on Databricks Operations/Platform SME who can independently administer, configure, monitor, troubleshoot, optimize, and support Databricks production environments, rather than a candidate focused primarily on Databricks data engineering/development.

Similar jobs