Haystack
← Back to Jobs
Remote
Technology

Principal Data Engineer

Saven TechnologiesSan Diego, CA🇺🇸United StatesPosted 30 Jul 2026

Quick Overview

Work Type
Remote
Level
Leader

Job Description

Hi,

 

Principal Data Engineer

Location: San Diego, CA

Way of Working: 3 days onsite, two days WFH

Interview process: 30 minutes

Interview Process: 1.) 30-minute interview 2.) 2-hour onsite 3.) offer

 

Principal Data Engineer

Join a pioneering research organization developing next-generation technology for high-precision industrial systems. Our engineering teams combine advanced instrumentation, sensing, controls, and physics-based modeling to address some of the most complex challenges in advanced manufacturing.

As our research organization continues to expand its use of data-driven engineering, machine learning, simulation, and physics-based modeling, a scalable and well-governed data ecosystem has become essential. High-quality, accessible, and connected data enables faster technology development, deeper system understanding, more effective trade studies, and better-informed technology and roadmap decisions.

In this role, you will help shape the data foundation that supports research and development activities across the organization. Working closely with lab owners and experimental, modeling, and ML scientists, you will build and improve data pipelines, integrate diverse data sources, and enable reliable access to research data at scale. You will also help establish practical architecture standards and best practices that ensure our data platform remains scalable, secure, maintainable, and aligned with the broader enterprise data landscape.

This role combines hands-on development with technical leadership in shaping the data foundation for advanced R&D. You will build, operate, and continuously improve data pipelines, integrating new data sources, improving reliability, and enabling scientists and engineers to use high-quality data at scale.

This is a Flex position with the potential to convert to a regular full-time position based on business needs, individual performance, and organizational priorities.

Responsibilities:

•              Define and evolve the data architecture strategy and standards for the research organization to enable data analytics and machine learning workflows.

•              Build and integrate data pipelines that connect research prototypes, experimental test benches, and simulation environments, ensuring data is discoverable, accessible, and reusable by scientists and engineers.

•              Establish data governance standards and best practices, including data lineage, access control, metadata management, security, and lifecycle policies.

•              Monitor and optimize data pipelines: implement quality controls and validation rules, track operational health, troubleshoot failures, and improve performance and cost efficiency.

•              Partner with teams across Research, Engineering, and IT to establish and align on a common data platform architecture.

•              Enable integration of physics-based models, AI capabilities, simulation workflows, and high-performance computing resources to support system-level understanding, analysis, and technology development.

•              Document platform architecture, design decisions, standards, and best practices, and communicate technical concepts effectively to both technical and non-technical stakeholders.

•              Work independently and collaboratively to deliver on objectives, whether exploring new data sources, building new capabilities, or characterizing existing system performance.

•              Be willing to work extended hours and second shift as needed.

•              Perform other duties as assigned or required.

Qualifications:

•              Bachelor''s or Master''s degree in Computer Science, Statistics, Math, Data Science, or a related field.

•              10+ years of relevant experience in data engineering, data architecture, or scientific/engineering data platforms.

•              Strong hands-on development experience in Python and modern data engineering tooling.

•              Proven experience building and operating scalable big data pipelines, analytics platforms, and data products that support data-intensive scientific and engineering workflows.

•              Experience with cloud and distributed data platforms such as Azure, Azure Databricks, Apache Spark, Kubernetes, and data lake architectures.

•              Solid understanding of data modeling, metadata management, data lineage, data quality, governance, security, and access control.

•              Experience supporting scientific or engineering workflows (e.g., simulation, HPC, instrumentation, or time-series sensor data).

•              Familiarity with AI/ML workflows and MLOps practices is a plus.

•              Strong communication and collaboration skills, with the ability to translate technical details into clear and actionable guidance.

 

 

Skills

MLOps
Machine Learning
Apache
Apache Spark
Azure
Databricks
Kubernetes
Python

Similar jobs