Haystack
← Back to Jobs
Technology

Lead Data Engineer

Vega Intellisoft Inc.Foster City, CA🇺🇸United StatesPosted 31 Jul 2026

Quick Overview

Work Type
On Site
Level
Mid Senior

Job Description

Hi Professionals,

Greetings from VegaIntellisoft

 

Role             : Lead Data Engineer

Location       : Foster City, CA (Onsite)
Duration       : Full time

Experience    : 10+ years needed

 

Hiring Data Engineering Lead with strong expertise in Databricks on AWS, MDM, and enterprise data integration. Lead the design and delivery of modern data platforms that enable trusted, governed, and scalable data consumption across business functions. Experience in cloud-based data engineering, middleware integrations, data governance, and enterprise-scale analytics solutions. Work closely with business, architecture, analytics, and engineering teams to drive data modernization initiatives. This role will be instrumental in enabling the enterprise data modernization journey. Establish a scalable and governed data foundation that supports advanced analytics, AI/ML initiatives, and business decision-making. Success in this role will directly improve data quality, consistency, and accessibility across critical business domains. The architecture and integration patterns defined by this role will serve as a foundation for future data and digital transformation programs.

 

Skills / Experience

  • 10+ years of experience in Data Engineering, Data Integration, or Data Platform delivery; hands-on experience with Databricks on AWS; Apache Spark, PySpark, Delta Lake, and Lakehouse architecture
  • Experience designing and implementing enterprise-scale data pipelines; Strong understanding of AWS services such as S3, Glue, Lambda, Redshift, IAM, and CloudWatch
  • Hands-on experience with MDM implementations and integrations; Experience with data quality, data governance, lineage, and master data management processes
  • Strong experience integrating enterprise systems using middleware platforms such as MuleSoft, Boomi, Kafka, or API-based integrations
  • Experience working with structured, semi-structured, and unstructured datasets; Strong SQL and Python development skills
  • Experience with Agile methodologies and DevOps practices; Experience leading distributed teams and managing stakeholder communications
  • Bachelor’s Degree or higher in Information Systems, Computer Science, or equivalent experience
  • Skills / Tech Stack Snapshot - Databricks on AWS, Apache Spark, PySpark, Delta Lake, AWS S3, Glue, Lambda, Redshift, MDM Platforms (Informatica MDM, Reltio, Profisee or equivalent), Data Integration & ETL/ELT Frameworks, Middleware Technologies (MuleSoft, Boomi, Kafka, API-led Integrations), Data Warehousing & Data Lake Architecture, Data Governance, MDM, Data Quality & Metadata Management, SQL, Python, Azure DevOps, Jira, Confluence, CI/CD, Agile Delivery

 

Job / Role Description

  • Lead the design and implementation of Databricks-based data platforms on AWS; Architect scalable Lakehouse solutions supporting enterprise analytics workloads
  • Design and develop complex ETL/ELT pipelines using Databricks, Spark, and Cloud-native services; Drive MDM strategy, implementation, and integration across business applications and data platforms
  • Define data integration patterns using APIs, middleware, event-driven architectures, and messaging frameworks; Establish data governance, metadata management, and data quality frameworks
  • Collaborate with business stakeholders to understand data requirements and translate them into technical solutions; Optimize data processing performance, scalability, and operational monitoring
  • Define CI/CD processes and deployment standards for data engineering assets; Mentor engineering teams and provide technical leadership throughout the project lifecycle
  • Support architecture reviews, solution design discussions, and technical decision-making; Ensure compliance with organizational standards, security requirements, and best practices
  • Facilitate architecture reviews, workshops, and stakeholder discussions; Ability to communicate complex technical concepts to business and executive stakeholders;
  • Stakeholder Management – Build strong relationships with business, IT, and external partners; Manage competing priorities and drive consensus among stakeholders; Demonstrate customer-centric and consultative engagement skills.
  • Leadership Skills – Lead cross-functional and geographically distributed teams; Mentor and guide engineers and junior architects; Influence technical decisions through collaboration
  • Problem Solving & Analytical Thinking – Identify root causes of complex data and integration challenges; Evaluate multiple solution options and recommend optimal approaches; Strong troubleshooting and performance optimization capabilities

 

Secondary Skills / Good to have

  • Experience with Snowflake or Microsoft Fabric; Exposure to AI/ML enablement using Databricks ML or AWS SageMaker
  • Experience with Unity Catalog and data governance frameworks; Data Mesh and Data Product concepts.
  • Experience with real-time streaming using Kafka or Kinesis; Informatica IDMC, Talend, or Azure Data Factory.
  • Knowledge of healthcare, life sciences, retail, manufacturing, or financial services domains.
  • Exposure to GenAI and enterprise AI adoption initiatives

Skills

SQL
AWS
ETL
Snowflake
Agile
Apache
Apache Spark
Azure
Confluence
Databricks
Jira
Kafka
Python
Redshift
Stakeholder Management
Unity

Similar jobs