Efficient Generative Adversarial Network Framework for Imbalanced Big Data Classification
DOI:
https://doi.org/10.64972/dea.2022.v1i1.2902d:16-28Keywords:
Generative Adversarial Network, Imbalanced Classification, Big Data, Minority Synthesis, Density RatioAbstract
Imbalanced Big Data classification is still challenging because the rare class has high operating value but provides weak statistical signals. Traditional re-sampling, cost-sensitive learning and ensemble strategies can improve the recall of the minority class, but they often fail to account for the density structure of the minority class and may add redundant or unsafe samples when the amount of data is very large. This paper proposes an efficient generative adversarial network framework for imbalanced big data classification, named EDR-GAN, which combines density-ratio-guided minority generation with prototype-preserving discrimination and scalable batch construction. The framework estimates local scarcity, constructs controlled conditional noise, and regularises the generated instances with respect to minority prototypes to expand informative regions rather than simply duplicating rare samples. Four representative imbalanced big data scenarios have been selected for the experimental design: transaction risk, network intrusion, equipment fault and medical screening. As shown in the results, the proposed framework enhances macro-F1, G-mean, AUC-PR and minority recall under imbalance ratios of 1:20 to 1:500, while reducing the training time compared to the heavier adversarial augmentation baseline. A reproducible engineering framework has been proposed to apply adversarial generation in large-scale classification systems for the detection of minority classes reliably without sacrificing deployment efficiency.
Downloads
Published
How to Cite
Issue
Section
License
Copyright (c) 2022 Dariusz Dębski, Henryk Homa

This work is licensed under a Creative Commons Attribution-NonCommercial-ShareAlike 4.0 International License.