Mayank Sethi
Mayank Sethi
Principal Database Engineer, Group 1001
Beyond "Throw an LLM at It": Tiered Computational Intelligence for Classifying 1.5 Million Columns of Enterprise Data
As enterprises consolidate decades of legacy systems into cloud data platforms, data governance faces a scale problem: how do you accurately classify millions of columns of sensitive data without an army of analysts — and without surrendering judgment entirely to a black-box model? This keynote presents the architecture and lessons from a production system that classified over 1.5 million columns across 62 databases in nine days, built and deployed by a single engineer. The system uses a tiered classification engine that applies deterministic, institution-aware rules first, native statistical classification second, and reserves large language model inference for only the truly ambiguous minority of columns — cutting cost by an order of magnitude while suppressing the over-classification that makes naive LLM tagging unusable at scale. Equally important is what surrounds the models: a human-in-the-loop review workflow that lets domain experts correct the engine without writing SQL, with corrections flowing back into production tags in about a minute. The talk argues for a design philosophy of "automation that serves experts, not replaces them," and offers a practical blueprint for regulated industries — finance, insurance, healthcare, where governance must be simultaneously fast, auditable, and trustworthy.