This contribution deals with the problem of text classification. The proposed approach is probabilistic and it is based on a mixture of a Dirichlet and Multinomial distributions. Our aim is to build a classifier able, not only to tale into account the words frequency, but also the latent topics contained within the available corpora. This new model, called sbDCM, allows us to insert directly the number of topics (known or unknown) that compound the document, without losing the 'burstiness' phenomenon and the classification performance.
Semantic based DCM models for text classification
CERCHIELLO, PAOLA
2012-01-01
Abstract
This contribution deals with the problem of text classification. The proposed approach is probabilistic and it is based on a mixture of a Dirichlet and Multinomial distributions. Our aim is to build a classifier able, not only to tale into account the words frequency, but also the latent topics contained within the available corpora. This new model, called sbDCM, allows us to insert directly the number of topics (known or unknown) that compound the document, without losing the 'burstiness' phenomenon and the classification performance.File in questo prodotto:
Non ci sono file associati a questo prodotto.
I documenti in IRIS sono protetti da copyright e tutti i diritti sono riservati, salvo diversa indicazione.