Graph-induced restricted Boltzmann machines for document modeling
dc.contributor.author | Nguyen, T. | |
dc.contributor.author | Tran, The Truyen | |
dc.contributor.author | Phung, D. | |
dc.contributor.author | Venkatesh, S. | |
dc.date.accessioned | 2017-01-30T15:23:59Z | |
dc.date.available | 2017-01-30T15:23:59Z | |
dc.date.created | 2016-03-17T19:30:18Z | |
dc.date.issued | 2016 | |
dc.identifier.citation | Nguyen, T. and Tran, T.T. and Phung, D. and Venkatesh, S. 2016. Graph-induced restricted Boltzmann machines for document modeling. Information Sciences. 328: pp. 60-75. | |
dc.identifier.uri | http://hdl.handle.net/20.500.11937/45888 | |
dc.identifier.doi | 10.1016/j.ins.2015.08.023 | |
dc.description.abstract |
© 2015 Elsevier Inc. All rights reserved. Discovering knowledge from unstructured texts is a central theme in data mining and machine learning. We focus on fast discovery of thematic structures from a corpus. Our approach is based on a versatile probabilistic formulation - the restricted Boltzmann machine (RBM) - where the underlying graphical model is an undirected bipartite graph. Inference is efficient - document representation can be computed with a single matrix projection, making RBMs suitable for massive text corpora available today. Standard RBMs, however, operate on bag-of-words assumption, ignoring the inherent underlying relational structures among words. This results in less coherent word thematic grouping. We introduce graph-based regularization schemes that exploit the linguistic structures, which in turn can be constructed from either corpus statistics or domain knowledge. We demonstrate that the proposed technique improves the group coherence, facilitates visualization, provides means for estimation of intrinsic dimensionality, reduces overfitting, and possibly leads to better classification accuracy. | |
dc.publisher | Elsevier Inc | |
dc.title | Graph-induced restricted Boltzmann machines for document modeling | |
dc.type | Journal Article | |
dcterms.source.volume | 328 | |
dcterms.source.startPage | 60 | |
dcterms.source.endPage | 75 | |
dcterms.source.issn | 0020-0255 | |
dcterms.source.title | Information Sciences | |
curtin.department | Multi-Sensor Proc & Content Analysis Institute | |
curtin.accessStatus | Fulltext not available |
Files in this item
Files | Size | Format | View |
---|---|---|---|
There are no files associated with this item. |