Haystack
← Back to Jobs
Technology
BL

Senior Data Engineer (ONLY FULLTIME)

Blackstraw LLCUnited States🇺🇸United StatesPosted Sep 16, 2026

Quick Overview

Seniority
Mid Senior
Work mode
Hybrid
Location
United States
Posted
Yesterday
SQLMLflowMachine LearningSnowflakeAirflowAzureComputer VisionDatabricksGenerative AIGitKafkaLLMPythonUnity

Job Description

Senior Data Engineer - (Pyspark, Databricks and Snowflake)

About the Company

Blackstraw.ai is an end-to-end technology services company specializing in Artificial Intelligence (AI) and Engineering solutions across Data Science, Data Engineering, LLM/GenAI and LLMOps. Founded in 2018, we help global enterprises across North America, Europe and Asia to build and operationalize AI systems that create measurable business impact. Our mission is to make AI adoption simpler, faster and scalable through a blend of deep domain expertise, reusable accelerators and proven engineering practices.

With a 600+ strong team of engineers, data scientists and AI specialists, we partner with organizations to deliver real-world outcomes in areas such as predictive analytics, computer vision, natural language processing and Generative AI. Headquartered in Florida (USA) with operations in Canada and India, Blackstraw.ai continues to empower global enterprises to unlock the true potential of AI.

Location: USA / Canada (Full‑time only)

MUST Experience: 10+ Years in Pyspark, Databricks and Snowflake

Role Summary

As a Senior Data Engineer, you will design, build, and optimize large‑scale data pipelines and analytics solutions. The role focuses on developing high‑performance data processing and transformation workflows using PySpark, orchestrating scalable compute workloads on Databricks, and implementing modern cloud‑native data architectures on Snowflake to support enterprise analytics and reporting needs.

Key Responsibilities

  • PySpark pipeline development — Design, develop, and maintain scalable data pipelines using PySpark and Databricks for batch and streaming workloads.

  • Databricks ELT workflows — Build and optimize ELT workflows using Databricks notebooks, Delta Lake, and distributed processing frameworks.

  • Snowflake data modeling — Develop high‑performance data models, transformations, and analytical datasets in Snowflake using SQL, Snowflake Streams, Tasks, and Warehouses.

  • Cloud data integration — Integrate structured and semi‑structured data sources using PySpark, Databricks Auto Loader, Snowpipe, Kafka, and cloud‑native ingestion tools.

  • Performance optimization — Optimize Databricks clusters, PySpark jobs, and Snowflake compute for performance, scalability, and cost efficiency.

  • Delta Lake & governance — Implement Delta Lake best practices including schema enforcement, DLT pipelines, data quality checks, and versioning.

  • Snowflake & Databricks administration — Manage and monitor Snowflake warehouses, Databricks jobs, Unity Catalog, and data access policies.

  • Cross‑functional collaboration — Work closely with data scientists, analysts, and business teams to deliver production‑ready data solutions.

  • Security & compliance — Implement data governance, RBAC, IAM, Purview, and enterprise‑grade security controls across Databricks and Snowflake.

  • Troubleshooting distributed systems — Diagnose and resolve issues across cloud data pipelines, PySpark jobs, Snowflake tasks, and streaming systems.

  • Architecture & documentation — Contribute to solution architecture, design reviews, and technical documentation for continuous improvement.

Required Skills & Experience

  • PySpark & Databricks expertise — Strong hands‑on experience with PySpark, Databricks, Delta Lake, and distributed data processing.

  • Snowflake proficiency — Solid experience with Snowflake including SQL, Warehouses, Streams, Tasks, and data modeling.

  • Cloud data engineering — Proficiency in Azure Data Factory, ADLS, Event Hub, and other cloud‑native ingestion and orchestration services.

  • Distributed systems knowledge — Strong understanding of distributed computing, scalable data pipelines, and cloud‑native architectures.

  • Advanced SQL & data modeling — Expertise in SQL, semi‑structured data handling, and dimensional/analytical data modeling.

  • Python programming — Strong programming skills in Python for data engineering, automation, and transformation logic.

  • CI/CD & DevOps — Knowledge of CI/CD pipelines, DevOps practices, and version control using Git.

Preferred Qualifications

  • Azure Certifications — Azure certifications such as DP‑203 or Azure Data Engineer Associate.

  • Orchestration Tools — Experience with Kafka, Airflow, or other cloud‑native orchestration and streaming tools.

  • ML Workflows — Exposure to machine learning workflows in Databricks, including model training, feature engineering, and MLflow.

Soft Skills

  • Strong analytical and problem‑solving abilities.

  • Clear communication and teamwork.

  • Ability to learn quickly and adapt to new technologies.

  • Ownership mindset and attention to detail.

Education

  • Bachelor’s degree in Computer Science, Engineering, IT, or equivalent experience.

Blackstraw provides equal employment opportunities to applicants and employees without regard to race, color, religion, age, sex, sexual orientation, gender identity/expression, national origin, marital status, protected veteran status, disability status, or any other basis as protected by federal, state, or local law

Similar jobs