Details of Research Outputs

Status已发表Published
TitleCombining tag and value similarity for data extraction and alignment
Creator
Date Issued2012
Source PublicationIEEE Transactions on Knowledge and Data Engineering
ISSN1041-4347
Volume24Issue:7Pages:1186-1200
Abstract

Web databases generate query result pages based on a user's query. Automatically extracting the data from these query result pages is very important for many applications, such as data integration, which need to cooperate with multiple web databases. We present a novel data extraction and alignment method called CTVS that combines both tag and value similarity. CTVS automatically extracts data from query result pages by first identifying and segmenting the query result records (QRRs) in the query result pages and then aligning the segmented QRRs into a table, in which the data values from the same attribute are put into the same column. Specifically, we propose new techniques to handle the case when the QRRs are not contiguous, which may be due to the presence of auxiliary information, such as a comment, recommendation or advertisement, and for handling any nested structure that may exist in the QRRs. We also design a new record alignment algorithm that aligns the attributes in a record, first pairwise and then holistically, by combining the tag and data value similarity information. Experimental results show that CTVS achieves high precision and outperforms existing state-of-the-art data extraction methods. © 2012 IEEE.

Keywordautomatic wrapper generation Data extraction data record alignment information integration
DOI10.1109/TKDE.2011.66
URLView source
Indexed BySCIE
Language英语English
WOS Research AreaComputer Science ; Engineering
WOS SubjectComputer Science, Artificial Intelligence ; Computer Science, Information Systems ; Engineering, Electrical & Electronic
WOS IDWOS:000304202000003
Scopus ID2-s2.0-84861736426
Citation statistics
Cited Times:18[WOS]   [WOS Record]     [Related Records in WOS]
Document TypeJournal article
Identifierhttp://repository.uic.edu.cn/handle/39GCC9TT/6599
CollectionFaculty of Science and Technology
Corresponding AuthorSu, Weifeng
Affiliation
1.Computer Science and Technology Program,BNU-HKBU United International College,Tangjiawan, Zhuhai,28, Jinfeng Road,China
2.Department of Computer Science,City University of Hong Kong,Kowloon,Tat Chee Avenue,Hong Kong
3.Department of Computer Science and Engineering,Hong Kong University of Science and Technology,Kowloon,Clear Water Bay,Hong Kong
4.Center for Speech and Language Technologies,Division of Technology Innovation and Development,Tsinghua National Laboratory for Information Science and Technology,Beijing,China
First Author AffilicationBeijing Normal-Hong Kong Baptist University
Corresponding Author AffilicationBeijing Normal-Hong Kong Baptist University
Recommended Citation
GB/T 7714
Su, Weifeng,Wang, Jiying,Lochovsky, Frederick H.et al. Combining tag and value similarity for data extraction and alignment[J]. IEEE Transactions on Knowledge and Data Engineering, 2012, 24(7): 1186-1200.
APA Su, Weifeng, Wang, Jiying, Lochovsky, Frederick H., & Liu, Yi. (2012). Combining tag and value similarity for data extraction and alignment. IEEE Transactions on Knowledge and Data Engineering, 24(7), 1186-1200.
MLA Su, Weifeng,et al."Combining tag and value similarity for data extraction and alignment". IEEE Transactions on Knowledge and Data Engineering 24.7(2012): 1186-1200.
Files in This Item:
There are no files associated with this item.
Related Services
Usage statistics
Google Scholar
Similar articles in Google Scholar
[Su, Weifeng]'s Articles
[Wang, Jiying]'s Articles
[Lochovsky, Frederick H.]'s Articles
Baidu academic
Similar articles in Baidu academic
[Su, Weifeng]'s Articles
[Wang, Jiying]'s Articles
[Lochovsky, Frederick H.]'s Articles
Bing Scholar
Similar articles in Bing Scholar
[Su, Weifeng]'s Articles
[Wang, Jiying]'s Articles
[Lochovsky, Frederick H.]'s Articles
Terms of Use
No data!
Social Bookmark/Share
All comments (0)
No comment.
 

Items in the repository are protected by copyright, with all rights reserved, unless otherwise indicated.