Extreme multi-label classification (XML) is becoming increasingly relevant in the era of big data. Yet, there is no method for effectively generating stratified partitions of XML datasets. Instead, researchers typically rely on provided test-train splits that, 1) aren't always representative of the entire dataset, and 2) are missing many of the labels. This can lead to poor generalization ability and unreliable performance estimates, as has been established in the binary and multiclass settings. As such, this paper presents a new and simple algorithm that can efficiently generate stratified partitions of XML datasets with millions of unique labels. We also examine the label distributions of prevailing benchmark splits, and investigate the issues that arise from using unrepresentative subsets of data for model development. The results highlight the difficulty of stratifying XML data, and demonstrate the importance of using stratified partitions for training and evaluation.
机构:
Aligarh Muslim Univ, Dept Stat & Operat Res, Aligarh, Uttar Pradesh, IndiaAligarh Muslim Univ, Dept Stat & Operat Res, Aligarh, Uttar Pradesh, India
Khowaja, Saman
Ghufran, Shazia
论文数: 0引用数: 0
h-index: 0
机构:
Aligarh Muslim Univ, Dept Stat & Operat Res, Aligarh, Uttar Pradesh, IndiaAligarh Muslim Univ, Dept Stat & Operat Res, Aligarh, Uttar Pradesh, India
Ghufran, Shazia
Ahsan, M. J.
论文数: 0引用数: 0
h-index: 0
机构:
Aligarh Muslim Univ, Dept Stat & Operat Res, Aligarh, Uttar Pradesh, IndiaAligarh Muslim Univ, Dept Stat & Operat Res, Aligarh, Uttar Pradesh, India
机构:
VIT AP Univ, Sch Adv Sci, Dept Math, Beside AP Secretariat, Amaravati, AP, IndiaVIT AP Univ, Sch Adv Sci, Dept Math, Beside AP Secretariat, Amaravati, AP, India
Triveni, G. R. V.
Danish, Faizan
论文数: 0引用数: 0
h-index: 0
机构:
VIT AP Univ, Sch Adv Sci, Dept Math, Beside AP Secretariat, Amaravati, AP, IndiaVIT AP Univ, Sch Adv Sci, Dept Math, Beside AP Secretariat, Amaravati, AP, India
机构:
Aligarh Muslim Univ, Dept Stat & Operat Res, Aligarh, Uttar Pradesh, IndiaAligarh Muslim Univ, Dept Stat & Operat Res, Aligarh, Uttar Pradesh, India
Khowaja, Saman
Ghufran, Shazia
论文数: 0引用数: 0
h-index: 0
机构:
Aligarh Muslim Univ, Dept Stat & Operat Res, Aligarh, Uttar Pradesh, IndiaAligarh Muslim Univ, Dept Stat & Operat Res, Aligarh, Uttar Pradesh, India
Ghufran, Shazia
Ahsan, M. J.
论文数: 0引用数: 0
h-index: 0
机构:
Aligarh Muslim Univ, Dept Stat & Operat Res, Aligarh, Uttar Pradesh, IndiaAligarh Muslim Univ, Dept Stat & Operat Res, Aligarh, Uttar Pradesh, India
机构:
VIT AP Univ, Sch Adv Sci, Dept Math, Beside AP Secretariat, Amaravati, AP, IndiaVIT AP Univ, Sch Adv Sci, Dept Math, Beside AP Secretariat, Amaravati, AP, India
Triveni, G. R. V.
Danish, Faizan
论文数: 0引用数: 0
h-index: 0
机构:
VIT AP Univ, Sch Adv Sci, Dept Math, Beside AP Secretariat, Amaravati, AP, IndiaVIT AP Univ, Sch Adv Sci, Dept Math, Beside AP Secretariat, Amaravati, AP, India