Haystack
← Back to Jobs
Technology
TE

Data Scientist

TechVirtue LLCUnited States🇺🇸United StatesPosted 11 Sept 2026

Quick Overview

Seniority
Mid Senior
Work mode
Hybrid
Location
United States
Posted
Yesterday
MLOpsMLflowMachine LearningDatabricksPythonUnity

Job Description

We are seeking a highly experienced Data Scientist / Databricks Architect to design and lead scalable data science and machine learning solutions using the Databricks Lakehouse platform. The ideal candidate will combine strong hands-on data science and machine learning expertise with the ability to define architecture, establish MLOps practices, and productionize ML solutions at enterprise scale.

The candidate should have deep experience with Databricks, Python, PySpark, MLflow, Delta Lake, Unity Catalog, machine learning, and cloud platforms, along with strong architecture and technical leadership capabilities.

Key Responsibilities

  • Design and architect enterprise-scale Data Science and Machine Learning solutions using Databricks.
  • Define end-to-end architecture for data ingestion, feature engineering, model development, deployment, monitoring, and retraining.
  • Develop and productionize machine learning models using Python, PySpark, and Databricks.
  • Establish best practices for ML lifecycle management, MLOps, model governance, and deployment.
  • Utilize MLflow for experiment tracking, model management, model registry, and deployment.
  • Design and implement scalable Delta Lake / Lakehouse architectures.
  • Implement Unity Catalog for data governance, security, access control, lineage, and discoverability.
  • Design Databricks workflows, jobs, clusters, notebooks, and production pipelines.
  • Develop scalable feature engineering and machine learning pipelines using Spark/PySpark.
  • Collaborate with Data Engineering, Cloud, DevOps, and Business teams to translate requirements into technical architecture.
  • Define architecture standards, reusable frameworks, design patterns, and engineering best practices.
  • Optimize Spark/Databricks workloads for performance, scalability, reliability, and cost.
  • Design solutions supporting batch and near-real-time machine learning workloads.
  • Implement model monitoring, performance tracking, drift detection, and automated retraining strategies.
  • Provide technical leadership and mentor Data Scientists, ML Engineers, and Data Engineers.
  • Evaluate emerging Databricks and AI/ML capabilities and recommend appropriate technologies.

Similar jobs