Model guided algorithm for mining unordered embedded subtrees

Hadzic, Fedja; Tan, H.; Dillon, Tharam S.

doi:10.3233/WIA-2010-0200

Access Status

Fulltext not available

Authors

Hadzic, Fedja

Tan, H.

Dillon, Tharam S.

Date

2010

Type

Journal Article

Metadata

Show full item record

Citation

Hadzic, Fedja and Tan, Henry and Dillon, Tharam S. 2010. Model guided algorithm for mining unordered embedded subtrees. Web Intelligence and Agent Systems. 8 (4): pp. 413-430.

Source Title

Web Intelligence and Agent Systems

DOI

10.3233/WIA-2010-0200

ISSN

15701263

School

Digital Ecosystems and Business Intelligence Institute (DEBII)

URI

http://hdl.handle.net/20.500.11937/37772

Collection

Curtin Research Publications

Abstract

Large amount of online information is or can be represented using semi-structured documents, such as XML. The information contained in an XML document can be effectively represented using a rooted ordered labeled tree. This has made the frequent pattern mining problem recast as the frequent subtree mining problem, which is a pre-requisite for association rule mining form tree-structured documents. Driven by different application needs a number of algorithms have been developed for mining of different subtree types under different support definitions. In this paper we present an algorithm for mining unordered embedded subtrees. It is an extension of our general tree model guided (TMG) candidate generation framework and the proposed U3 algorithm considers all support definitions, namely, transaction-based, occurrence-match and hybrid support. A number of experiments are presented on synthetic and real world data sets. The results demonstrate the flexibility of our general TMG framework as well as its efficiency when compared to the existing state-of-the-art approach.