Haystack
← Back to Jobs
Technology
TS

Staff Data Engineer - Python / Google Cloud Platform

The Search Solutions, LLCUnited States🇺🇸United StatesPosted 31 Aug 2026

Quick Overview

Seniority
Mid Senior
Work mode
Hybrid
Location
United States
Posted
23 hours ago
DjangoFastAPIMicroservicesNeo4jSQLAWSAirflowApacheBigQueryGoogle CloudPython

Job Description

The Staff Data Engineer is a hands-on role and will be the primary architect and technical lead for the data infrastructure powering our next-generation Agentic AI products. Acting as a hands-on leader, you are responsible for the team’s overall delivery, translating complex product requirements into actionable technical tasks for a small engineering squad. You will design the multi-modal data stores (Vector and Graph) that serve as the "Active Memory" for autonomous agents while remaining deeply embedded in the codebase to drive execution.

Key Responsibilities

Technical Execution: Lead the technical delivery by translating high-level product roadmaps into actionable development cycles. You will own the task breakdown and manage the workflow to ensure high-quality output from the team

Strategic Data Architecture: Architect and directly implement multi-modal data pipelines that process structured parts catalogs and unstructured sources (PDFs, Word, PNG/SVG diagrams) into specialized Vector and Graph data stores.

Knowledge Layer Development: Design and optimize the storage of embeddings in Vector Databases (e.g., Pinecone, ChromaDB, Vertex AI Search) and Graph Databases (e.g., FalkorDB, Neo4j) to enable multi-step agentic reasoning across disparate data sources.

Agent-Driven Development: Deeply integrate autonomous coding agents into your daily workflow to plan, generate, and refactor data infrastructure and microservices.

Evaluation Pipelines: Collaborate with AI Engineers to build "Gold Dataset" pipelines used for the automated verification, retrieval quality measurement, and "confidence scoring" of AI outputs.

Microservices: Develop and maintain scalable data services and RESTful APIs using Python (FastAPI/Django) to provide structured, validated data.

Cloud Operations: Deploy and monitor high-scale data workloads on Google Cloud Platform (Vertex AI, BigQuery), ensuring system reliability, security, and cost-effectiveness

Skills and Experience

10+ years of Software/Data Engineering experience, with a proven history of leading technical teams.

Proven track record of designing and building production-grade AI data systems.

Technical Stack: Advanced Python (FastAPI, Pydantic, SQLAlchemy) and SQL mastery for building scalable microservices.

AI: Hands-on experience with Vector Databases (Pinecone, ChromaDB), RAG pipelines, and GraphRAG patterns.

Data Tooling: Deep experience with Prefect (preferred) or Apache Airflow or Cloud Composer, BigQuery, and DataProc.

Cloud Infrastructure: Experienced in working with cloud platforms (Google Cloud Platform, AWS) and deploying data workloads and pipelines at scale.

Coding Agent: Demonstrated proficiency in using coding agents to accelerate the SDLC and plan and code complex engineering tasks.

Similar jobs