Haystack
← Back to Jobs
Technology
SA

Sr Data Engineer

SG Analytics Inc.Los Angeles, CA🇺🇸United StatesPosted Sep 16, 2026

Quick Overview

Seniority
Mid Senior
Work mode
Hybrid
Location
Los Angeles, CA, United States
Posted
Yesterday
SQLAWSFlinkMLOpsMLflowAirflowApacheDatabricksGitUnity

Job Description

Role: Sr Data Engineer (Full Time/FTE)
Location: LA
Experience: 8+ Years
Work Model: Hybrid (3 days in a week)

Primary Skill: Delta Lake, ML workflows, PySpark, and AWS, Databricks,Delta Lake,Notebooks, and SQL, distributed processing frameworks,  EMR, Dataproc, open-source Spark 

Job Description
       The Data Engineering team is seeking a Senior Data Engineer to help design, build, and scale the modern data platform that powers analytics, data science, and data products. In this role, you’ll collaborate closely with Data Product, Data Science, Analytics, and Engineering teams to deliver reliable, high-impact data solutions used by hundreds of internal users. We're looking for engineers who enjoy hands-on development, take ownership of production systems, and influence implementation through collaboration and technical expertise.

Key Responsibilities
Design and implement lakehouse architecture using Delta Lake, including medallion pipeline patterns (Bronze/Silver/Gold), schema enforcement, and time travel

Build and operate batch and real-time ingestion pipelines leveraging Databricks Auto Loader, Structured Streaming, and Change Data Capture patterns

Implement data governance and security using Unity Catalog, RBAC, and compliance-driven practices for sensitive environments

Optimize performance and manage costs through FinOps strategies, including cluster sizing, workload tagging, Spark tuning, and Photon acceleration

Design, implement, and maintain CI/CD pipelines and orchestration workflows using Databricks Workflows, Delta Live Tables, and tools such as Airflow

Collaborate with Data Science teams on ML workflows, including MLflow, feature store integration, and model lifecycle management

Ensure data quality, observability, and lineage across media-specific datasets such as streaming logs, ad impressions, and audience metrics

Provide technical mentorship through code reviews, pairing and knowledge sharing

Qualifications
Bachelor’s degree in Computer Science, Data Engineering, or equivalent practical experience

5+ years of experience building production-grade data pipelines in cloud environments using Spark-based platforms (e.g., Databricks, EMR, Dataproc, open-source Spark)

Expertise in PySpark, SQL, and Spark-based data processing, with experience operating pipelines at scale in production

Hands-on experience building batch or streaming production data pipelines using distributed processing frameworks (e.g., Spark, Flink) and query engines such as Presto

Proficiency with orchestration tools such as Apache Airflow or Dagster, with hands-on experience in CI/CD, monitoring, alerting, and data quality for production systems

Experience working with modern data architectures, including event-driven and distributed systems

Proficiency with Git and collaborative development workflows

Build and operate batch and real-time ingestion pipelines using Spark based batch and streaming patterns (e.g., Structured Streaming, CDC), with experience on Databricks or comparable platforms

Solid understanding of infrastructure, networking, and data security fundamentals


Preferred Qualifications
Experience building Lakehouse platforms and medallion pipelines in Databricks

Familiarity with Unity Catalog, data governance, and compliance frameworks (e.g., PCI)

Hands-on experience with CI/CD pipelines, orchestration tools, and infrastructure-as-code

Experience with Lakeflow Spark Declarative Pipelines (SDP), MLflow, feature stores, and MLOps practices

Background in media and entertainment data (e.g., video metadata, ad tech, audience analytics)

Experience building data platforms within the media industry, with a strong understanding of audience analytics

Experience working with large-scale analytical datasets (e.g., event logs, clickstream data, audience metrics)

Comfortable using AI-assisted development tools (e.g., ChatGPT)

Databricks or cloud certifications (e.g., Databricks Certified Data Engineer)

Next Steps & Details Needed:

       If this role aligns with your career goals, please reply to this email with your updated resume and the quick details below:


Full Name:
Work Authorization:
Validity of Visa (if any expiry):
Current Location:
Current Company:
LinkedIn Profile URL:
Email ID : 
Phone Number:
Total Years in Data Engineering:
Open to Full-Time and Hybrid :
Notice Period / Availability:

 

Similar jobs