A Rebalancing Framework for Classification of Imbalanced Medical Appointment No-show Data,Journal of Data and Information Science

当前位置： X-MOL 学术 › Journal of Data and Information Science › 论文详情

Our official English website, www.x-mol.net, welcomes your feedback! (Note: you will need to create a separate account there.)

A Rebalancing Framework for Classification of Imbalanced Medical Appointment No-show Data
Journal of Data and Information Science ( IF 1.5 ) Pub Date : 2021-01-27 , DOI: 10.2478/jdis-2021-0011
Ulagapriya Krishnan ₁ , Pushpa Sangar ₁

Affiliation

Abstract

Purpose

This paper aims to improve the classification performance when the data is imbalanced by applying different sampling techniques available in Machine Learning.

Design/methodology/approach

The medical appointment no-show dataset is imbalanced, and when classification algorithms are applied directly to the dataset, it is biased towards the majority class, ignoring the minority class. To avoid this issue, multiple sampling techniques such as Random Over Sampling (ROS), Random Under Sampling (RUS), Synthetic Minority Oversampling TEchnique (SMOTE), ADAptive SYNthetic Sampling (ADASYN), Edited Nearest Neighbor (ENN), and Condensed Nearest Neighbor (CNN) are applied in order to make the dataset balanced. The performance is assessed by the Decision Tree classifier with the listed sampling techniques and the best performance is identified.

Findings

This study focuses on the comparison of the performance metrics of various sampling methods widely used. It is revealed that, compared to other techniques, the Recall is high when ENN is applied CNN and ADASYN have performed equally well on the Imbalanced data.

Research limitations

The testing was carried out with limited dataset and needs to be tested with a larger dataset.

Practical implications

This framework will be useful whenever the data is imbalanced in real world scenarios, which ultimately improves the performance.

Originality/value

This paper uses the rebalancing framework on medical appointment no-show dataset to predict the no-shows and removes the bias towards minority class.

中文翻译：

不平衡医疗预约未出现数据分类的再平衡框架

摘要

目的

本文旨在通过应用机器学习中可用的不同采样技术来提高数据不平衡时的分类性能。

设计/方法/方法

医疗预约未出现数据集是不平衡的，并且当分类算法直接应用于数据集时，它偏向多数类别，而忽略了少数类别。为避免此问题，采用了多种采样技术，例如随机过采样（ROS），随机欠采样（RUS），合成少数采样技术（SMOTE），自适应合成采样（ADASYN），最近邻编辑（ENN）和最近邻压缩（CNN）以便使数据集平衡。决策树分类器使用列出的采样技术评估性能，并确定最佳性能。