Conversational Data Engineer
Experience: 8+ Years
Work Model: Remote / Hybrid
Role Overview
We are seeking a hands-on Synthetic Data & RAG / Knowledge Pipeline Engineer to build high-quality synthetic datasets and scalable knowledge pipelines supporting conversational AI, voice, and chat solutions.
This role focuses on creating realistic, privacy-safe data for AI configuration, testing, evaluation, and demonstrations, while also owning the ingestion, indexing, retrieval, and evaluation processes required for Retrieval-Augmented Generation (RAG) solutions.
The ideal candidate combines experience in data engineering, applied NLP, conversational AI data, synthetic data generation, and RAG/knowledge engineering.
Key Responsibilities
Analyze business domains to understand intents, entities, knowledge topics, document types, languages, conversational patterns, and edge cases.
Design and generate high-quality synthetic conversational datasets for voice and chat use cases.
Generate training and evaluation dialogues, customer/agent profiles, FAQs, knowledge-base articles, and structured entity data.
Build and maintain RAG and knowledge ingestion pipelines for enterprise conversational AI solutions.
Prepare, clean, transform, chunk, and organize documents and other knowledge sources for retrieval.
Load and manage data across appropriate cloud storage, data, search, and knowledge services.
Implement indexing and retrieval processes to ensure relevant information can be accurately retrieved by AI applications.
Evaluate retrieval quality, coverage, diversity, realism, and overall data quality.
Build parameterized and reusable data-generation and ingestion pipelines rather than customer-specific manual processes.
Implement and document knowledge-index refresh processes as source information changes.
Ensure synthetic datasets and source corpora meet privacy and data-isolation requirements, including validation for potential PII exposure.
Document generation approaches, parameters, assumptions, quality checks, and known limitations.
Collaborate with conversational AI, platform, and DevOps engineering teams to ensure datasets and knowledge indexes work correctly within deployed solutions.
Required Qualifications
4+ years of experience in data engineering, applied NLP, conversational AI data, knowledge engineering, or related areas.
Hands-on experience building or supporting RAG / knowledge retrieval pipelines.
Understanding of document ingestion, preprocessing, chunking, indexing, retrieval, and evaluation.
Experience working with conversational datasets such as transcripts, FAQs, training dialogues, intents, entities, or knowledge articles.
Strong understanding of data quality, retrieval quality, and privacy/PII considerations.
Experience building automated, reusable data-processing or ingestion pipelines.
Strong analytical and problem-solving skills.
Preferred Qualifications
Hands-on experience with LLM-assisted synthetic data generation in production or enterprise implementation environments.
Experience generating synthetic conversations, evaluation datasets, knowledge articles, or similar AI datasets.
Familiarity with GCP and AWS data services, including technologies such as BigQuery, Amazon S3, and document/knowledge stores.
Experience with vector search, semantic retrieval, embeddings, or enterprise search technologies.
Experience evaluating RAG systems using retrieval and data-quality metrics.
Experience with multilingual synthetic-data generation or evaluation.
Familiarity with conversational AI, contact-center AI, voice bots, or enterprise chat applications.