Quick Overview
Job Description
• Create high-quality, customer-specific synthetic data.
• Own RAG / knowledge pipelines for each deployment of CCAI, voice, and chat.
• Configure, ground, demonstrate, and validate deployments without using real customer PII.
• Design generation and ingestion pipelines and load data into correct Google Cloud Platform and AWS services.
• Document synthetic dataset for each customer engagement, covering channels in scope.
• Ensure each in-scope customer has a working RAG / knowledge pipeline: corpus prepared, indexed, retrievable, and evaluated.
• Validate data and retrieval quality for configuration, evaluation, and stakeholder demos.
• Ensure safe for isolation and compliance expectations.
• Parameterize and repeat generation and indexing, not a one-off manual copy-paste per customer.
• Analyze each customer’s domain: intents, entities, knowledge topics, document types, languages, tone, and edge cases.
• Generate synthetic conversation transcripts for voice and chat, plus CCAI training/evaluation dialogues.
• Generate supporting content: customer/agent profiles, knowledge-base articles, FAQs, and structured entity values.
• Schedule and document index refresh processes when customer knowledge changes.
• Use appropriate techniques while documenting parameters and limitations.
• Validate realism, coverage, diversity, and absence of residual real-world PII in synthetic data and source corpora.
• Maintain reusable generators, ingestion jobs, and quality checklists that can be parameterized per customer.
• Partner
• The Conversational Platform Specialist is responsible for loading data and indexes to drive the deployed experience.
• They must work with DevOps to automate and isolate pipeline jobs, stores, and secrets per customer.
• Required qualifications include 4+ years of experience in data engineering, conversation design operations, applied NLP data work, or knowledge-pipeline engineering.
• They should have a working knowledge of how conversational platforms consume training, FAQ, transcript, and retrieval-grounded knowledge data.
• Strong judgment on synthetic-data quality, retrieval quality, and privacy safety is necessary.
• Preferred qualifications include experience with LLM-assisted synthetic data generation in a production or implementation setting, familiarity with BigQuery, S3, and document stores used as knowledge sources, and multilingual data generation or evaluation experience.
• Work experience of 7-10 years is required.
Similar jobs
- CI
AWS Data Engineer with AI Skills - Charlotte, NC (Hybrid 2 Days/Week Onsite)
NewCuboid IT Solutions
Charlotte, NC🇺🇸On-site19 hours agoSQLAWSETL+8Technology - GI
Data Engineer
NewGeorgia IT
Boston, MA🇺🇸Hybrid19 hours agoSQLSQL ServerETL+7Technology - NI
Hadoop and PySpark Data Engineer
NewNityo Infotech Corporation
Jersey City, NJ🇺🇸Hybrid19 hours agoSQLShellAgile+4Technology - SD
AI Data Engineer
NewStanley David and Associates
United States🇺🇸Hybrid19 hours agoAWSETLSnowflake+2Technology - DT
Sr. Data Engineer
NewDonato Technologies Inc
United States🇺🇸Hybrid19 hours agoOracleSQLETL+3Technology - LS
kafka engineer
NewLincoln Softtech LLC
Irvine, CA🇺🇸Hybrid19 hours agoSAFeAgileEngineering