Posts

Showing posts with the label Synthetic Data for RAG

Garbage In, Liability Out: Cleaning Unstructured Data with AI

 In traditional data warehousing, dirty data usually means null values or duplicates. We fix that with simple, deterministic rules. In Generative AI, dirty data is far more dangerous. It shows up as semantic noise —and it can change what the model believes is true. Consider a financial-services chatbot powered by RAG (Retrieval-Augmented Generation). It ingests thousands of marketing PDFs. Every page contains a legal footer: “This document contains forward-looking statements that are not guarantees of future performance.” Now a user asks, “What is the projected growth?” The retriever pulls the footer. The model reads the legal language, misinterprets it, and responds: “The company guarantees future performance.” That isn’t a UX bug. It’s a liability. This is the new reality of data hygiene. You can’t just dump raw PDFs into a vector database and hope for the best. You must clean them first. And because the problem is semantic, not structural, standard code isn’t enough. You need an...

The 'Cold Start' Fix: Generating Synthetic Golden Sets with Unity Catalog

There is a moment in nearly every RAG project that I call the Evaluation Deadlock . The engineering team has built a chatbot. It works. They’re ready to test it. Then someone asks the obvious question: “What should we test it against?” The room goes quiet. To measure quality, you need a Golden Dataset—at least 100 realistic user questions paired with accurate, ground-truth answers. But early on, you don’t have users yet. Which means you don’t have questions. And your Subject Matter Experts—the senior lawyers, engineers, or doctors who could write those answers—are far too busy billing $500 an hour to spend days in Excel creating test cases. This is the deadlock: You can’t deploy without testing. You can’t test without data. The solution isn’t to hire more humans. It’s to build an Evaluation Factory. Here’s how to use Databricks Unity Catalog and synthetic data generation to bootstrap a high-quality test suite overnight—without consuming a single hour of SME time.