HARMONY: Efficiently Mining the Best Rules for Classification
2004-09-16
Loading...
View/Download File
Persistent link to this item
Statistics
View StatisticsJournal Title
Journal ISSN
Volume Title
Title
HARMONY: Efficiently Mining the Best Rules for Classification
Alternative title
Authors
Published Date
2004-09-16
Publisher
Type
Report
Abstract
Many studies have shown that rule-based classification algorithms perform well in classifying categorical and sparse high-dimensional databases. However, a fundamental limitation with many rule-based classifiers is that they find the classification rules in a coarse-grained manner. They usually use heuristic methods to prune the search space, and select the rules based on the sequential database covering paradigm. Thus, the so-mined rules may not be the globally best rules for some instances in the training database. To make worse, these algorithms fail to fully exploit some more effective search space pruning methods in order to scale to large databases.
In this paper we propose a new classifier, HARMONY, which directly mines the final set of classification rules. HARMONY uses an instance-centric rule-generation approach in the sense that it can assure for each training instance, one of the highest-confidence rules covering this instance is included in the result set, which helps a lot in achieving high classification accuracy. By introducing several novel search strategies and pruning methods into the traditional frequent itemset mining framework, HARMONY also has high efficiency and good scalability. Our thorough performance study with some large text and categorical databases has shown that HARMONY outperforms many well-known classifiers in terms of both accuracy and efficiency, and scales well w.r.t. the database size.
Keywords
Description
Related to
Replaces
License
Series/Report Number
Technical Report; 04-032
Funding information
Isbn identifier
Doi identifier
Previously Published Citation
Other identifiers
Suggested citation
Wang, Jianyong; Karypis, George. (2004). HARMONY: Efficiently Mining the Best Rules for Classification. Retrieved from the University Digital Conservancy, https://hdl.handle.net/11299/215626.
Content distributed via the University Digital Conservancy may be subject to additional license and use restrictions applied by the depositor. By using these files, users agree to the Terms of Use. Materials in the UDC may contain content that is disturbing and/or harmful. For more information, please see our statement on harmful content in digital repositories.