Universidad de Salamanca
Clasificación automática de información en portales web mediante técnicas de clustering
Abstract
[EN]The expression, "Information Retrieval" (Information Retrieval), refers to automatic processing is carried out in order to respond to a need for information. It includes some aspects of the representation, storage and organization of information and on the other aspects of efficiency in the production of results as a result of consultations. These provide the user with valuable information that is relevant, not only data as far as possible classified and weighted as to their degree of usefulness. There are various classification algorithms have been used. They tend to operate according to a set of premises, which in many cases will be measured or measurable, which will result in different models of information retrieval. Classic models such as Boolean, the vector or probabilistic. Alternative to the classical models, such as finite sets, boolean extended, generalized vector space, that of latent semantic indexing, the neural network, the network of inferences or network of beliefs. We have aimed to give an overview of all of them and a classification. Moreover clustering techniques are techniques of data analysis in which observations are applied according to their similarity. Its fields of application are most diverse: business, microeconomics, GIS, bioinformatics, genomics, image segmentation, natural language processing and a long list that includes aspects we want to address as the classification of documents in Recovery Information. Some have sought to apply the rules of clustering for automated sorting large amounts of information that usually handle directories of many web sites, looking for shared libraries documents in groups, so that they can subsequently be applied to other practical problems such as the grouping of documents obtained from web searches, or viewing of directories. This has been necessary to analyze the various document clustering techniques to analyze their methods and determine which one best fits the classification of documents from web sites and model a process by identifying and characterizing the different phases that combines models of recovery technical information clustering approaches. This addresses a topic of great interest to the user of information technology and communication as the improvement in the location of content to the growing flood of data and its temporality. On the other hand seeks to provide new forms of web portals to present information to complement existing ones.
Author and committee
dc:creator, dc:contributor.*- Author
-
- Alvarez García, Juan Carlos
Subjects
dc:subject × 9Identifiers
dc:identifier.*- Identifier
- hdl:10366/76382
- OAI identifier oai:identifier
- oai:gredos.usal.es:10366/76382