Pratik Pravin Mahajan

Pratik Pravin Mahajan

Pratik Pravin Mahajan

Sr. Data Analyst, JPMorgan Chase & Co.

Title of Talk:

Data Quality as a Machine Learning Problem: Detecting Silent Data Corruption Before It Reaches the Executive Dashboard


Abstract

The majority of enterprise data failures go unnoticed by an alert. Pipelines run without a hitch, dashboards are refreshed on time, and numbers are being fed to leaders but quietly corrupted, this is known as silent data corruption. Business rules that are written by hand only reveal the errors that teams already knew about the biggest ones don't catch them. This session flips the approach to enterprise data quality from a manual checklist to a machine learning approach. The talk will try to give a real-world recipe for how to build an automated framework for self-monitoring data pipelines, based on experience building self-monitoring pipelines for large-scale cloud data migrations, cutting manual ETL processing time by 40% and maintaining automated reporting that is up and running in global regions for the financial services industry with a 93% uptime. They include statistical profiling of data sets, distribution-shift and anomaly detection to raise the red flags of unexpected data corruption, learned validation rules vs. static checks, scoring a data set before it's published, and when human review should still be in the chain in regulated environments. The audience will learn to become aware of their own pipeline's silent failure modes, how to use anomaly detection tools to monitor data quality, when to use machine-learned data checks instead of rule-based validation, and how to create a dataset trust score to preserve the credibility of enterprise analytics and decision making.


Brief Profile
I am a Senior Data Analyst with over 5 years of experience working within the Business Intelligence/Data Engineering/Data Analytics space in Financial Services and Regulated industries. My focus is on data quality validation, ETL automation, and dashboard governance, from leading the data validation effort for an enterprise Oracle-to-Databricks data migration to the cloud to developing automated reporting and forecasting solutions that are deployed by global stakeholders. I am a strong Python and SQL developer and experienced user of Tableau, Sigma, Databricks and Alteryx with professional writing and research interests in data governance, responsible AI and analytics reliability for financial institutions. I have an M.S. degree in Data Analytics and Information Technology from Rutgers University and an MBA degree in Finance in progress. In addition to my primary job, I provide mentorship to junior analysts, provide analytics to assist senior executives in decision making and am a peer reviewer for data analytics and business intelligence.