Back to results

University of North Texas

The Cluster Hypothesis: A Visual/Statistical Analysis

Abstract

dc:description

By allowing judgments based on a small number of exemplar documents to be applied to a larger number of unexamined documents, clustered presentation of search results represents an intuitively attractive possibility for reducing the cognitive resource demands on human users of information retrieval systems. However, clustered presentation of search results is sensible only to the extent that naturally occurring similarity relationships among documents correspond to topically coherent clusters. The Cluster Hypothesis posits just such a systematic relationship between document similarity and topical relevance. To date, experimental validation of the Cluster Hypothesis has proved problematic, with collection-specific results both supporting and failing to support this fundamental theoretical postulate. The present study consists of two computational information visualization experiments, representing a two-tiered test of the Cluster Hypothesis under adverse conditions. Both experiments rely on multidimensionally scaled representations of interdocument similarity matrices. Experiment 1 is a term-reduction condition, in which descriptive titles are extracted from Associated Press news stories drawn from the TREC information retrieval test collection. The clustering behavior of these titles is compared to the behavior of the corresponding full text via statistical analysis of the visual characteristics of a two-dimensional similarity map. Experiment 2 is a dimensionality reduction condition, in which inter-item similarity coefficients for full text documents are scaled into a single dimension and then rendered as a two-dimensional visualization; the clustering behavior of relevant documents within these unidimensionally scaled representations is examined via visual and statistical methods. Taken as a whole, results of both experiments lend strong though not unqualified support to the Cluster Hypothesis. In Experiment 1, semantically meaningful 6.6-word document surrogates systematically conform to the predictions of the Cluster Hypothesis. In Experiment 2, the majority of the unidimensionally scaled datasets exhibit a marked nonuniformity of distribution of relevant documents, further supporting the Cluster Hypothesis. Results of the two experiments are profoundly question-specific. Post hoc analyses suggest that it may be possible to predict the success of clustered searching based on the lexical characteristics of users' natural-language expression of their information need.

Degree

thesis:*
Grantor dc:publisher
University of North Texas
Year dc:date
2000

Author and committee

dc:creator, dc:contributor.*
Author dc:creator
  • Sullivan, Terry
Contributors dc:contributor
  • Norris, Cathleen
  • Cleveland, Donald B., 1935-
  • Pavur, Robert J.
  • Schamber, Linda
  • Young, Jon I.

Subjects

dc:subject × 7

Rights

dc:rights
Statement dc:rights
  • Use restricted to UNT Community
  • Copyright
  • Sullivan, Terry
  • Copyright is held by the author, unless otherwise noted. All rights reserved.
Language dc:language
English

Identifiers

dc:identifier.*
Identifier
oclc: 47233220
untcat: b2302287
https://digital.library.unt.edu/ark:/67531/metadc2444/
ark: ark:/67531/metadc2444
OAI identifier oai:identifier
info:ark/67531/metadc2444

Chain of custody

source
Harvested from
University of North Texas
Base URL
digital.library.unt.edu/oai/
Last updated
2026-07-24
Source record
OAI-PMH GetRecord
citation

Sullivan, Terry. The Cluster Hypothesis: A Visual/Statistical Analysis. University of North Texas, 2000. https://doi.org/10.12794/metadc2444