Haystack
← Back to Jobs
Technology
IF

Conversational Data Engineer - AILab-GENAI-Google Cloud Platform.

IT First SourceUnited States🇺🇸United StatesPosted 9 Sept 2026

Quick Overview

Seniority
Mid Senior
Work mode
Hybrid
Location
United States
Posted
19 hours ago
AWSNLPBigQueryGoogle CloudLLM

Job Description

    •    Create high-quality, customer-specific synthetic data.
    •    Own RAG / knowledge pipelines for each deployment of CCAI, voice, and chat.
    •    Configure, ground, demonstrate, and validate deployments without using real customer PII.
    •    Design generation and ingestion pipelines and load data into correct Google Cloud Platform and AWS services.
    •    Document synthetic dataset for each customer engagement, covering channels in scope.
    •    Ensure each in-scope customer has a working RAG / knowledge pipeline: corpus prepared, indexed, retrievable, and evaluated.
    •    Validate data and retrieval quality for configuration, evaluation, and stakeholder demos.
    •    Ensure safe for isolation and compliance expectations.
    •    Parameterize and repeat generation and indexing, not a one-off manual copy-paste per customer.
    •    Analyze each customer’s domain: intents, entities, knowledge topics, document types, languages, tone, and edge cases.
    •    Generate synthetic conversation transcripts for voice and chat, plus CCAI training/evaluation dialogues.
    •    Generate supporting content: customer/agent profiles, knowledge-base articles, FAQs, and structured entity values.
    •    Schedule and document index refresh processes when customer knowledge changes.
    •    Use appropriate techniques while documenting parameters and limitations.
    •    Validate realism, coverage, diversity, and absence of residual real-world PII in synthetic data and source corpora.
    •    Maintain reusable generators, ingestion jobs, and quality checklists that can be parameterized per customer.
    •    Partner
    •    The Conversational Platform Specialist is responsible for loading data and indexes to drive the deployed experience.
    •    They must work with DevOps to automate and isolate pipeline jobs, stores, and secrets per customer.
    •    Required qualifications include 4+ years of experience in data engineering, conversation design operations, applied NLP data work, or knowledge-pipeline engineering.
    •    They should have a working knowledge of how conversational platforms consume training, FAQ, transcript, and retrieval-grounded knowledge data.
    •    Strong judgment on synthetic-data quality, retrieval quality, and privacy safety is necessary.
    •    Preferred qualifications include experience with LLM-assisted synthetic data generation in a production or implementation setting, familiarity with BigQuery, S3, and document stores used as knowledge sources, and multilingual data generation or evaluation experience.
    •    Work experience of 7-10 years is required.

Similar jobs