Quick Overview
Job Description
The Staff Data Engineer is a hands-on role and will be the primary architect and technical lead for the data infrastructure powering our next-generation Agentic AI products. Acting as a hands-on leader, you are responsible for the team’s overall delivery, translating complex product requirements into actionable technical tasks for a small engineering squad. You will design the multi-modal data stores (Vector and Graph) that serve as the "Active Memory" for autonomous agents while remaining deeply embedded in the codebase to drive execution.
Key Responsibilities
Technical Execution: Lead the technical delivery by translating high-level product roadmaps into actionable development cycles. You will own the task breakdown and manage the workflow to ensure high-quality output from the team
Strategic Data Architecture: Architect and directly implement multi-modal data pipelines that process structured parts catalogs and unstructured sources (PDFs, Word, PNG/SVG diagrams) into specialized Vector and Graph data stores.
Knowledge Layer Development: Design and optimize the storage of embeddings in Vector Databases (e.g., Pinecone, ChromaDB, Vertex AI Search) and Graph Databases (e.g., FalkorDB, Neo4j) to enable multi-step agentic reasoning across disparate data sources.
Agent-Driven Development: Deeply integrate autonomous coding agents into your daily workflow to plan, generate, and refactor data infrastructure and microservices.
Evaluation Pipelines: Collaborate with AI Engineers to build "Gold Dataset" pipelines used for the automated verification, retrieval quality measurement, and "confidence scoring" of AI outputs.
Microservices: Develop and maintain scalable data services and RESTful APIs using Python (FastAPI/Django) to provide structured, validated data.
Cloud Operations: Deploy and monitor high-scale data workloads on Google Cloud Platform (Vertex AI, BigQuery), ensuring system reliability, security, and cost-effectiveness
Skills and Experience
10+ years of Software/Data Engineering experience, with a proven history of leading technical teams.
Proven track record of designing and building production-grade AI data systems.
Technical Stack: Advanced Python (FastAPI, Pydantic, SQLAlchemy) and SQL mastery for building scalable microservices.
AI: Hands-on experience with Vector Databases (Pinecone, ChromaDB), RAG pipelines, and GraphRAG patterns.
Data Tooling: Deep experience with Prefect (preferred) or Apache Airflow or Cloud Composer, BigQuery, and DataProc.
Cloud Infrastructure: Experienced in working with cloud platforms (Google Cloud Platform, AWS) and deploying data workloads and pipelines at scale.
Coding Agent: Demonstrated proficiency in using coding agents to accelerate the SDLC and plan and code complex engineering tasks.
Similar jobs
- BT
Master Data Engineer | Configuration Management
BETA Technologies
South Burlington🇺🇸3 days agoCADCATIAERP+3Technology - ME
Senior Data Engineer (Agentic AI / RAG Platform) - Onsite - FORD- Locals Only
NewMeganSoft
Dearborn, MI🇺🇸On-site23 hours agoSQLAzurePythonTechnology - AS
Data Engineer 3
NewApex Systems
Atlanta, GA🇺🇸$55 - $60/hrOn-site23 hours agoSQLSQL ServerETL+5Technology - BA
Senior Help Desk Data Engineer
NewBooz Allen Hamilton
Winchester, VA🇺🇸$61.9k - $141k/yrOn-site23 hours agoMySQLOracleRust+16Technology - LM
Azure Data Engineer with AI Experience : Remote
NewLightning Minds Inc.
United States🇺🇸Remote23 hours agoSQLAzurePostgreSQL+2Technology - RI
Data Engineer
NewRITWIK Infotech Inc
Dallas, TX🇺🇸Hybrid23 hours agoSQLLookerData Pipeline+4Technology