Haystack
← Back to Jobs
Remote
Technology
ML

Data Scientist (Big Data Engineer)

Masterapp LabsAustin, TX🇺🇸United StatesPosted 3 Sept 2026

Quick Overview

Seniority
Mid Senior
Work mode
Remote
Location
Austin, TX, United States
Posted
18 hours ago
SQLETLEncryptionMLflowScikit-learnAgileApacheApache SparkAzureDatabricksLLMPythonRTensorFlowUnity

Job Description

Job Title: Data Scientist (Big Data Engineer)
Location: Austin, TX, (Telework). Mostly Remote sometimes weekly need to visit the office for in person meetings
Position Type: Contract
Interview Mode: Webcam and In-person both

Duties include:

  • Designing and developing scalable data pipelines
  • Implementing ETL/ELT workflows
  • Optimizing Spark jobs
  • Integrating with Azure Data Factory
  • Automating deployments
  • Collaborating with cross-functional teams
  • Ensuring data quality, governance, and security
  • Build AI accelerators and reusable components to boost development teams’ productivity.
II.  CANDIDATE SKILLS AND QUALIFICATIONS
Minimum Requirements:
Candidates that do not meet or exceed the minimum stated requirements (skills/experience) will be displayed to customers but may not be chosen for this opportunity.
Years
Required/Preferred
Experience
8
Required
Implement ETL/ELT workflows for both structured and unstructured data
8
Required
Collaborate with cross-functional teams including data scientists, analysts, and stakeholders
8
Required
Design and maintain data models, schemas, and database structures to support analytical and operational use cases
8
Required
Implement data validation and quality checks to ensure accuracy and consistency
8
Required
Contribute to data governance initiatives, including metadata management, data lineage, and data cataloging
8
Required
Implement data security measures, including encryption, access controls, and auditing; ensure compliance with regulations and best practices
8
Required
Working in agile, multicultural environments
4
Required
Automate deployments using CI/CD tools
4
Required
Evaluate and implement appropriate data storage solutions, including data lakes (Azure Data Lake Storage) and data warehouses
4
Required
Proficiency in Python and R programming languages
4
Required
Strong SQL querying and data manipulation skills
4
Required
Experience with Azure cloud platform
4
Required
Experience with DevOps, CI/CD pipelines, and version control systems
4
Required
Strong troubleshooting and debugging capabilities
3
Required
Design and develop scalable data pipelines using Apache Spark on Databricks
3
Required
Optimize Spark jobs for performance and cost efficiency
3
Required
Integrate Databricks solutions with cloud services (Azure Data Factory)
3
Required
Ensure data quality, governance, and security using Unity Catalog or Delta Lake
3
Required
Deep understanding of Apache Spark architecture, RDDs, DataFrames, and Spark SQL
3
Required
Hands-on experience with Databricks notebooks, clusters, jobs, and Delta Lake
2
Required
Build AI accelerators and reusable components to boost development teams’ productivity.
2
Required
Implement Gen AI / LLM application and utilities (Agentic AI, Harness / Prompt engineering, RAG)
2
Required
Implement AI application governance, observability and evaluation framework
2
Required
Design & Implement vector stores for knowledge bases and agent memory
2
Required
Integrate AI applications / utils / tools with enterprise traceability and SIEM tools
2
Required
Knowledge of ML libraries (MLflow, Scikit-learn, TensorFlow)
1
Preferred
Databricks Certified Associate Developer for Apache Spark
1
Preferred
Azure Data Engineer Associate

Similar jobs