A comprehensive framework for de-duplication: Acute kidney failure (AKF) case study
Chomchanok Yawana et al.
What the paper says
<b>Objectives:</b> Addressing data duplication is one of the most important issues in electronic health record (EHR) processing since the nature of data collection in the field. It does not only affect the data quality in healthcare management, but also the reliability in the downstream analyses. In this paper, we propose a comprehensive data de-duplication framework tailored for medical databases to tackle data duplication for a kidney disease identification, Acute Kidney Failure (AKF). <b>Methods:</b> The proposed work begins with the data joining from various sources, basic data de-duplication which automatically removes the dirty texts, medical note-event extraction since the data could be sources for further de-duplication, NLP data de-duplication based on a pre-trained model, data mapping for integration, unrelated data and outlier elimination, and eventually data imputation by a clustered based imputer. <b>Results:</b> We illustrated our de-duplication framework on MIMIC-III database both on the de-duplication task and the classification task based on AKF. The experiments demonstrated that the proposed work could achieve up to 99.59% accuracy or 23% higher than the traditional method and could achieve a high classification accuracy at 86 % and the F1-score at 0.87, which outperformed the traditional method, and the original dataset without any modification. <b>Conclusion:</b> These results demonstrated that the framework can potentially address the data duplication issue in healthcare effectively.
Evidence weight
Balanced mode · F 0.40 / M 0.15 / V 0.05 / R 0.40
| F · citation impact | 0.50 × 0.4 = 0.20 |
| M · momentum | 0.50 × 0.15 = 0.07 |
| V · venue signal | 0.50 × 0.05 = 0.03 |
| R · text relevance † | 0.50 × 0.4 = 0.20 |
† Text relevance is estimated at 0.50 on the detail page — for your query’s actual relevance score, open this paper from a search result.