A Two-Stage Classifier with Reject Option for Text Categorisation

Publication Type:

Conference Paper

Source:

5th Int. Workshop on Statistical Techniques in Pattern Recognition (SPR 2004), Springer, Volume 3138, Lisbon, Portugal, p.771-779 (2004)

Keywords:

document categorisation; text categorisation; classification reliability; reject option; rej00; doc01; doc00

Abstract:

Abstract. In this paper, we investigate the usefulness of the reject option in text categorisation systems. The reject option is introduced by allowing a text classifier to withhold the decision of assigning or not a document to any subset of categories, for which the decision is considered not sufficiently reliable. To automatically handle rejections, a two-stage classifier architecture is used, in which documents rejected at the first stage are automatically classified at the second stage, so that no rejections eventually remain. The performance improvement achievable by using the reject option is assessed on a real text categorisation task, using the well known Reuters data set.

AttachmentSize
Fumera_SPR04.pdf151.61 KB