Posts

Showing posts with the label Databricks GenAI

The 'Cold Start' Fix: Generating Synthetic Golden Sets with Unity Catalog

There is a moment in nearly every RAG project that I call the Evaluation Deadlock . The engineering team has built a chatbot. It works. They’re ready to test it. Then someone asks the obvious question: “What should we test it against?” The room goes quiet. To measure quality, you need a Golden Dataset—at least 100 realistic user questions paired with accurate, ground-truth answers. But early on, you don’t have users yet. Which means you don’t have questions. And your Subject Matter Experts—the senior lawyers, engineers, or doctors who could write those answers—are far too busy billing $500 an hour to spend days in Excel creating test cases. This is the deadlock: You can’t deploy without testing. You can’t test without data. The solution isn’t to hire more humans. It’s to build an Evaluation Factory. Here’s how to use Databricks Unity Catalog and synthetic data generation to bootstrap a high-quality test suite overnight—without consuming a single hour of SME time.