Theses and Dissertations
Issuing Body
Mississippi State University
Advisor
Hodges, Julia E.
Committee Member
Jamil, Hasan M.
Committee Member
Bridges, Susan M.
Date of Degree
5-7-2005
Document Type
Graduate Thesis - Open Access
Major
Computer Science
Degree Name
Master of Science
College
James Worth Bagley College of Engineering
Department
Department of Computer Science and Engineering
Abstract
While progress has been made in querying digital information contained in XML and HTML documents, success in retrieving information from the so called "hidden Web" (data behind Web forms) has been modest. There has been a nascent trend of developing autonomous tools for extracting information from the hidden Web. Automatic tools for ontology generation, wrapper generation, Weborm querying, response gathering, etc., have been reported in recent research. This thesis presents a system called Chameleon for automatic querying of and response gathering from the hidden Web. The approach to response gathering is based on automatic table structure identification, since most information repositories of the hidden Web are structured databases, and so the information returned in response to a query will have regularities. Information extraction from the identified record structures is performed based on domain knowledge corresponding to the domain specified in a query. So called "domain plug-ins" are used to make the dynamically generated wrappers domain-specific, rather than conventionally used document-specific.
URI
https://hdl.handle.net/11668/19642
Recommended Citation
Chouvarine, Philippe, "Autonomous Consolidation of Heterogeneous Record-Structured HTML Data in Chameleon" (2005). Theses and Dissertations. 831.
https://scholarsjunction.msstate.edu/td/831