Pratik Pravin Mahajan
Pratik Pravin Mahajan
Sr. Data Analyst, JPMorgan Chase & Co.
Data Quality as a Machine Learning Problem: Detecting Silent Data Corruption Before It Reaches the Executive Dashboard
The majority of enterprise data failures go unnoticed by an alert. Pipelines run without a hitch, dashboards are refreshed on time, and numbers are being fed to leaders but quietly corrupted, this is known as silent data corruption. Business rules that are written by hand only reveal the errors that teams already knew about the biggest ones don't catch them. This session flips the approach to enterprise data quality from a manual checklist to a machine learning approach. The talk will try to give a real-world recipe for how to build an automated framework for self-monitoring data pipelines, based on experience building self-monitoring pipelines for large-scale cloud data migrations, cutting manual ETL processing time by 40% and maintaining automated reporting that is up and running in global regions for the financial services industry with a 93% uptime. They include statistical profiling of data sets, distribution-shift and anomaly detection to raise the red flags of unexpected data corruption, learned validation rules vs. static checks, scoring a data set before it's published, and when human review should still be in the chain in regulated environments. The audience will learn to become aware of their own pipeline's silent failure modes, how to use anomaly detection tools to monitor data quality, when to use machine-learned data checks instead of rule-based validation, and how to create a dataset trust score to preserve the credibility of enterprise analytics and decision making.