TGT: A Novel Adversarial Guided Oversampling Technique for Handling Imbalanced Datasets

Loading...
Thumbnail Image

Date

1/19/2021

Journal Title

Journal ISSN

Volume Title

Type

Article

Publisher

Elsevier

Series Info

Egyptian Informatics Journal;

Abstract

With the volume of data increasing exponentially, there is a growing interest in helping people to benefit from their data regardless of its poor quality. One of the major data quality problems is the imbalanced distribution of different categories existing in the data. Such problem would affect the performance of any possible of analysis and mining on the data. For instance, data with an imbalanced distribution has a negative effect on the performance achieved by most traditional classification techniques. This paper proposes TGT (Train Generate Test), a novel oversampling technique for handling imbalanced datasets problem. Using different learning strategies, TGT guarantees that the generated synthetic samples reside in minority regions. TGT showed a high improvement in performance of different classification techniques when was experimented with five imbalanced datasets of different types. 2021 Production and hosting by Elsevier B.V. on behalf of Faculty of Computers and Information, Cairo University. This is an open access article under the CC BY-NC-ND license (http://creativecommons.org/ licenses/by-nc-nd/4.0/).

Description

Keywords

university, Imbalance, Oversampling, Classification

Citation