TGT: A Novel Adversarial Guided Oversampling Technique for Handling Imbalanced Datasets
Loading...
Date
1/19/2021
Journal Title
Journal ISSN
Volume Title
Type
Article
Publisher
Elsevier
Series Info
Egyptian Informatics Journal;
Scientific Journal Rankings
Abstract
With the volume of data increasing exponentially, there is a growing interest in helping people to benefit from their data regardless of its poor quality. One of the major data quality problems is the imbalanced distribution of different categories existing in the data. Such problem would affect the performance of any possible of analysis and mining on the data. For instance, data with an imbalanced distribution has a negative effect on the performance achieved by most traditional classification techniques. This paper proposes TGT (Train Generate Test), a novel oversampling technique for handling imbalanced datasets problem. Using different learning strategies, TGT guarantees that the generated synthetic samples reside in minority regions. TGT showed a high improvement in performance of different classification techniques when was experimented with five imbalanced datasets of different types. 2021 Production and hosting by Elsevier B.V. on behalf of Faculty of Computers and Information, Cairo University. This is an open access article under the CC BY-NC-ND license (http://creativecommons.org/ licenses/by-nc-nd/4.0/).
Description
Keywords
university, Imbalance, Oversampling, Classification