Research Highlights

Towards Building Trustworthy and Practical Deep Clustering Net

Update as of 3 August 2026


Deep clustering, an important branch of unsupervised learning, has achieved significant progress in recent years. However, most existing methods assume balanced or near-balanced data distributions, which contrast with the long-tailed characteristics commonly observed in real-world datasets. This discrepancy often results in biased learning and degraded clustering performance. Although various long-tailed learning approaches have been developed, they typically depend on labeled data to guide class-specific adjustments, limiting their applicability in unsupervised settings.

To address this challenge, MiniClustering[1] is proposed as a mini-cluster guided framework for long-tailed deep clustering. The approach employs a specialized clustering head that partitions data into many fine-grained groups, referred to as mini-clusters. These mini-clusters enable the estimation of data imbalance without label information. Based on this estimation, adaptive weights are generated and incorporated into the self-training loss, effectively re-weighting gradients and mitigating model bias across clusters.

Extensive evaluation of benchmark datasets with varying imbalance ratios demonstrates the effectiveness of the proposed approach. The framework is compatible with existing unsupervised representation learning methods and facilitates the extension of label-dependent long-tailed techniques to fully unsupervised clustering scenarios.

 


Team Members:

  1. PI: Dr. LIU Hui, Yam Pak Charitable Foundation School of Computing and Information Sciences, Saint Francis University
  2. Dr. HOU Junhui, Department of Computer Science, City University of Hong Kong

Reference
[1] Zhixin Li, Yuheng Jia, Guanliang Chen, Hui Liu, Junhui Hou, Mini-cluster Guided Long-tailed Deep Clustering, The Fourteenth International Conference on Learning Representations, 2026.




Reference no.: UGC/FDS11/E03/24