Query expansion using local and global document analysis

Jinxi XuW. Bruce Croft

article1996SIGIR1,332 citationsTest of Time Award

Introduces local context analysis, a query expansion technique combining global word-context features with locally retrieved passages to outperform traditional corpus-wide and local feedback retrieval methods.

Listen

Information retrieval systems frequently fail because users submit short queries using different vocabulary than the authors of relevant documents. Automatic query expansion solves this word mismatch problem by enriching queries with related terms. However, traditional techniques face severe trade-offs: analyzing an entire text repository is computationally expensive and introduces off-topic terms, while standard local feedback based on initial search results performs unpredictably when early matches contain non-relevant documents.

The article evaluates existing query expansion techniques and introduces a hybrid method called local context analysis, which applies structural and context-aware term selection to top-ranked passages rather than entire documents or full collections. The authors conducted comparative experiments across three major benchmark collections—TREC-3, TREC-4, and the legal WEST database—evaluating overall precision, recall, and query-by-query reliability across hundreds of thousands of documents.

The analysis demonstrates that local context analysis consistently outperforms both global collection-level analysis and traditional local feedback. On large collections, local context analysis improved average retrieval precision by roughly 23% to 24% over unexpanded baseline searches, compared to modest gains of 3% to 8% for global analysis and 14% to 21% for local feedback. Furthermore, local context analysis proved substantially more robust on difficult queries: on the TREC-4 set, it improved 38 of 49 queries and harmed only 11, whereas standard local feedback degraded 21 queries—frequently destroying performance on queries that started with low initial precision. Finally, local context analysis showed low sensitivity to parameter tuning, maintaining high effectiveness across a broad window of 30 to 300 retrieved passages.

These findings indicate that search systems can achieve significant gains in retrieval accuracy and reliability without the massive computational and storage overhead required to build corpus-wide concept databases. By extracting noun phrases from focused 300-word passages and enforcing term co-occurrence constraints across all original query words, local context analysis provides an efficient mechanism suitable for interactive, real-time search environments.

Organizations implementing automated search expansion should adopt passage-level local context analysis over whole-document local feedback, especially when serving domains where individual search failure carries high operational risk. When applying expansion to specialized databases with high baseline accuracy and fewer relevant documents, systems should downweight expansion terms to prevent query drift. Future technical work should focus on developing adaptive algorithms that automatically calibrate passage counts and term weights for each query rather than relying on static system parameters.

Cover for Query expansion using local and global document analysis

Abstract

Automatic query expansion has long been suggested as a technique for dealing with the fundamental issue of word mismatch in information retrieval. A number of approaches to expansion have been studied and, more recently, attention has focused on techniques that analyze the corpus to discover word relationships (global techniques) and those that analyze documents retrieved by the initial query ( local feedback). In this paper, we compare the effectiveness of these approaches and show that, although global analysis has some advantages, local analysis is generally more effective. We also show that using global analysis techniques, such as word context and phrase structure, on the local set of documents produces results that are both more effective and more predictable than simple local feedback.

Table of Contents

  • 1 Introduction
  • 2 Global Analysis
  • 3 Local Analysis
  • 3.1 Local Feedback
  • 3.2 Local Context Analysis
  • 4 Experiments
  • 4.1 Collections and Query Sets
  • 4.2 Local Context Analysis
  • 5 Local Text Analysis vs Global Analysis
  • 6 Local Text Analysis vs Local Feedback
  • 7 Conclusion and Future Work
  • 8 Acknowledgements
  • References

Knowls

  1. Knowl 1 — Concept Ranking Formula in Local Context Analysis

    equation

    In Local Context Analysis (LCA), candidate concepts (defined as noun groups consisting of single nouns, two adjacent nouns, or three adjacent nouns) extracted from the top nn retrieved passages are ranked by their belief score bel(Q,c)bel(Q, c) with respect to query QQ:

    bel(Q,c)=∏ti∈Q(δ+log⁡(af(c,ti))⋅idfclog⁡(n))idfibel(Q, c) = \prod_{t_i \in Q} \left( \delta + \frac{\log(af(c, t_i)) \cdot idf_c}{\log(n)} \right)^{idf_i}

    where the constituent components are defined as:

    af(c,ti)=∑j=1nftij⋅fcjaf(c, t_i) = \sum_{j=1}^{n} ft_{ij} \cdot fc_j

    idfi=max⁡(1.0,log⁡10(N/Ni)5.0)idf_i = \max\left(1.0, \frac{\log_{10}(N / N_i)}{5.0}\right)

    idfc=max⁡(1.0,log⁡10(N/Nc)5.0)idf_c = \max\left(1.0, \frac{\log_{10}(N / N_c)}{5.0}\right)

    Variables and parameters:

    • QQ: the input query, composed of terms ti∈Qt_i \in Q.
    • cc: a candidate concept (noun group).
    • nn: the number of top-ranked passages retrieved by the initial query (n=100n = 100 by default).
    • pjp_j: the jj-th passage among the top nn retrieved passages (j∈{1,…,n}j \in \{1, \dots, n\}).
    • ftijft_{ij}: the frequency of query term tit_i in passage pjp_j.
    • fcjfc_j: the frequency of concept cc in passage pjp_j.
    • af(c,ti)af(c, t_i): the co-occurrence frequency of concept cc and query term tit_i across the top nn passages.
    • NN: the total number of passages in the entire document collection.
    • NiN_i: the number of passages in the collection that contain query term tit_i.
    • NcN_c: the number of passages in the collection that contain candidate concept cc.
    • idfiidf_i: the inverse passage frequency weight of query term tit_i, emphasizing rare query terms.
    • idfcidf_c: the inverse passage frequency weight of candidate concept cc, penalizing concepts that occur ubiquitously throughout the collection.
    • δ\delta: an additive smoothing constant set to 0.10.1 to avoid a zero belief value when a concept co-occurs with only a subset of query terms.

    The multiplicative product over all ti∈Qt_i \in Q ensures that selected concepts co-occur across all query terms rather than heavily co-occurring with only one non-discriminative term.

  2. Knowl 2 — Query Expansion Formulation and Concept Weighting in Local Context Analysis

    equation

    Once candidate concepts are scored and ranked via Local Context Analysis, the top mm concepts c1,c2,…,cmc_1, c_2, \dots, c_m are added to the original query QQ to form an expanded query QnewQ_{new} using the INQUERY weighted sum operator (#WSUM)(\#\text{WSUM}):

    Qnew=#WSUM(1.0  1.0  Q  w  Q′)Q_{new} = \#\text{WSUM}(1.0 \; 1.0 \; Q \; w \; Q')

    Q′=#WSUM(1.0  w1  c1  w2  c2  …  wm  cm)Q' = \#\text{WSUM}(1.0 \; w_1 \; c_1 \; w_2 \; c_2 \; \dots \; w_m \; c_m)

    where:

    • QQ: the original user query.
    • Q′Q': the auxiliary expansion query formed by the top mm selected concepts.
    • mm: the total number of expansion concepts added, set to m=70m = 70.
    • cic_i: the ii-th ranked concept according to bel(Q,c)bel(Q, c).
    • wiw_i: the weight assigned to concept cic_i, defined by a linearly decreasing rank-based schedule: w_i = 1.0 - 0.9 \cdot \frac{i}{70} \quad \text{for } i \in {1, 2, \dots, 70}$$
    • ww: the weight multiplier assigned to the auxiliary expansion query Q′Q' relative to the original query QQ. By default, w=2.0w = 2.0. For high-baseline queries where the original formulation is already very precise (such as legal text search on the WEST collection), ww is reduced to 1.01.0 (a 50% downweighting of expansion concepts).
    • #WSUM\#\text{WSUM}: the INQUERY weighted sum operator, which computes a normalized weighted average of belief scores across its arguments: #WSUM(w0  w1  A1  …  wk  Ak)=∑j=1kwj⋅score(Aj)∑j=1kwj\#\text{WSUM}(w_0 \; w_1 \; A_1 \; \dots \; w_k \; A_k) = \frac{\sum_{j=1}^k w_j \cdot \text{score}(A_j)}{\sum_{j=1}^k w_j}.
  3. Knowl 3 — Local Context Analysis Query Expansion Algorithm

    algorithm

    The Local Context Analysis (LCA) query expansion procedure operates in two phases: an offline collection frequency pass and an online query expansion and retrieval pass.

    Input: Query QQ, collection text database partitioned into fixed passages of length 300 words, number of passages nn (default 100), number of expansion concepts mm (default 70), auxiliary weight ww (default 2.0).
    Output: Expanded query QnewQ_{new} and final ranked document list.
    // Step 1: Passage Retrieval
    Retrieve the top nn ranked passages p1,p2,…,pnp_1, p_2, \dots, p_n for query QQ using baseline IR retrieval.
    // Step 2: Concept Extraction
    Extract all candidate concepts CC from the retrieved passages {p1,…,pn}\{p_1, \dots, p_n\}, where each concept c∈Cc \in C is a noun group (a single noun, 2 adjacent nouns, or 3 adjacent nouns).
    // Step 3: Concept Scoring
    for each concept c∈Cc \in C do
        idfc←max⁡(1.0,log⁡10(N/Nc)/5.0)idf_c \leftarrow \max(1.0, \log_{10}(N / N_c) / 5.0)
        score←1.0score \leftarrow 1.0
        for each term ti∈Qt_i \in Q do
            af(c,ti)←∑j=1nftij⋅fcjaf(c, t_i) \leftarrow \sum_{j=1}^n ft_{ij} \cdot fc_j
            idfi←max⁡(1.0,log⁡10(N/Ni)/5.0)idf_i \leftarrow \max(1.0, \log_{10}(N / N_i) / 5.0)
            term_factor←(0.1+(log⁡(af(c,ti))⋅idfc)/log⁡(n))idfiterm\_factor \leftarrow (0.1 + (\log(af(c, t_i)) \cdot idf_c) / \log(n))^{idf_i}
            score←score⋅term_factorscore \leftarrow score \cdot term\_factor
        end for
        bel(Q,c)←scorebel(Q, c) \leftarrow score
    end for
    // Step 4: Expansion Query Construction
    Sort concepts in CC descending by bel(Q,c)bel(Q, c) and select the top mm concepts c1,c2,…,cmc_1, c_2, \dots, c_m.
    for i←1i \leftarrow 1 to mm do
        wi←1.0−0.9⋅(i/m)w_i \leftarrow 1.0 - 0.9 \cdot (i / m)
    end for
    Construct auxiliary query Q′←#WSUM(1.0,w1,c1,w2,c2,…,wm,cm)Q' \leftarrow \#\text{WSUM}(1.0, w_1, c_1, w_2, c_2, \dots, w_m, c_m)
    Construct expanded query Qnew←#WSUM(1.0,1.0,Q,w,Q′)Q_{new} \leftarrow \#\text{WSUM}(1.0, 1.0, Q, w, Q')
    // Step 5: Final Retrieval
    Execute QnewQ_{new} against the full document collection and return ranked documents.

    Offline, the passage collection frequencies (Ni,NcN_i, N_c) are pre-computed in a single indexing pass taking approximately 3 hours on an Alpha workstation for a 2 GB corpus. Online expansion of a query across 100 passages requires only several seconds of CPU time.

  4. Knowl 4 — Passage-Level Retrieval Units in Local Context Analysis

    model/method

    Local Context Analysis extracts candidate expansion concepts from fixed-size text passages (windows of 300 words) rather than full documents. This design choice provides two distinct advantages:

    1. Elimination of Spurious Co-occurrences: Documents in large full-text collections can be very long and address multiple distinct subtopics. In a full document, co-occurrence between a candidate concept at the beginning of a document and a query term at the end often reflects accidental co-presence rather than semantic relatedness. Restricting co-occurrence tracking to 300-word passages ensures that terms and concepts occur within a shared, tight semantic context.
    2. Computational Efficiency: Using passages avoids parsing, indexing, and scoring unretrieved and irrelevant portions of long documents, reducing both disk I/O and online query expansion computation to a few seconds.
  5. Knowl 5 — Multiplicative vs. Additive Term Scoring for Query Expansion Selection

    model/method

    Global analysis techniques (such as Phrasefinder in INQUERY) use additive belief functions to rank candidate concepts. A significant failure mode of additive scoring is that candidate concepts that co-occur heavily with several non-discriminative query terms can dominate the ranking, even if they never co-occur with the primary semantic subject of the query. For example, for the query "As a result of DNA testing, are more defendants being absolved or convicted of crimes", additive scoring selects "oil spill" because it frequently co-occurs with "result", "test", "defendant", "absolve", and "crime", despite having zero co-occurrence with the central term "DNA".

    Local Context Analysis prevents this distortion by using a multiplicative product across all query terms:

    ∏ti∈Q(δ+log⁡(af(c,ti))⋅idfclog⁡(n))idfi\prod_{t_i \in Q} \left( \delta + \frac{\log(af(c, t_i)) \cdot idf_c}{\log(n)} \right)^{idf_i}

    The product structure forces a candidate concept to demonstrate co-occurrence across every query term (weighted by its term rarity idfiidf_i) to achieve a high score. Furthermore, because LCA limits candidate extraction to the top-ranked local passages, it does not require hard stop-filters on high-frequency concepts, allowing valid, high-frequency expansion concepts (such as "China" and "Iraq" for a query on nuclear proliferation treaties) that global analysis systems discard as too frequent.

  6. Knowl 6 — Retrieval Effectiveness of Local Context Analysis vs. Global and Local Feedback on TREC-4

    data/table

    On the TREC-4 benchmark (49 queries, topics 202–250, Tipster disks 2 and 3, 567,529 documents), Local Context Analysis (LCA with 100 passages) substantially outperforms both the unexpanded baseline, global document analysis (Phrasefinder adding 30 concepts), and standard local feedback (adding top 50 terms and 10 phrases from the top 10 documents via Rocchio α:β:γ=1:1:0\alpha:\beta:\gamma = 1:1:0).

    Recall Base Phrasefinder Local Feedback (10 doc) LCA (100 passages)
    0 71.0 68.6 (-3.3%) 68.4 (-3.6%) 73.2 (+3.2%)
    10 49.3 48.6 (-1.6%) 52.8 (+7.0%) 57.1 (+15.7%)
    20 40.4 40.0 (-1.0%) 43.2 (+7.0%) 46.8 (+16.0%)
    30 33.3 33.9 (+1.8%) 36.0 (+8.0%) 39.9 (+19.8%)
    40 27.3 28.0 (+2.5%) 29.8 (+9.2%) 35.3 (+29.1%)
    50 21.6 23.9 (+10.3%) 24.5 (+13.2%) 29.9 (+38.4%)
    60 14.8 18.8 (+27.1%) 19.7 (+33.4%) 23.6 (+59.8%)
    70 9.5 11.8 (+24.7%) 14.8 (+56.9%) 17.9 (+89.1%)
    80 6.2 8.1 (+31.0%) 10.8 (+74.7%) 11.8 (+91.0%)
    90 3.1 4.2 (+33.6%) 6.4 (+104.6%) 5.7 (+80.2%)
    100 0.4 0.6 (+24.0%) 0.9 (+93.3%) 0.8 (+88.2%)
    Average 25.2 26.0 (+3.4%) 27.9 (+11.0%) 31.1 (+23.5%)

    While Phrasefinder degrades precision at low recall levels (0–20% recall) and yields an overall average precision improvement of only +3.4%, LCA improves precision across all recall levels, achieving a +23.5% gain in 11-point average precision (increasing from 25.2% to 31.1%). LCA also outperforms local feedback at every recall level from 0% to 80%.

  7. Knowl 7 — Retrieval Effectiveness of Local Context Analysis vs. Global and Local Feedback on TREC-3

    data/table

    On the TREC-3 benchmark (50 queries, topics 151–200, Tipster disks 1 and 2, 741,856 documents), Local Context Analysis (100 passages) achieves higher average precision than both Phrasefinder (global analysis) and standard Rocchio local feedback on top 10 documents.

    Recall Base Phrasefinder Local Feedback (10 doc) LCA (100 passages)
    0 82.2 79.4 (-3.3%) 82.5 (+0.4%) 87.0 (+5.9%)
    10 57.3 60.1 (+4.8%) 64.9 (+13.3%) 65.5 (+14.3%)
    20 46.2 50.4 (+9.1%) 56.1 (+21.5%) 57.2 (+23.8%)
    30 39.1 43.3 (+10.7%) 48.3 (+23.5%) 48.4 (+23.8%)
    40 32.7 36.9 (+12.8%) 41.6 (+26.9%) 42.7 (+30.4%)
    50 27.5 31.8 (+15.9%) 36.8 (+34.1%) 37.9 (+38.0%)
    60 22.6 26.1 (+15.1%) 30.9 (+36.7%) 31.5 (+39.3%)
    70 18.0 20.6 (+14.0%) 25.2 (+40.0%) 25.6 (+42.1%)
    80 13.3 15.8 (+18.6%) 19.4 (+45.7%) 19.4 (+45.7%)
    90 7.9 9.4 (+18.7%) 11.5 (+44.3%) 11.7 (+47.3%)
    100 0.5 0.8 (+60.9%) 1.2 (+143.5%) 1.4 (+177.0%)
    Average 31.6 34.1 (+7.8%) 38.0 (+20.5%) 38.9 (+23.3%)

    LCA achieves a 23.3% relative improvement over the baseline at 100 passages (and peaks at +24.4% with 200 passages, reaching 39.3% average precision). Global analysis via Phrasefinder hurts precision at 0% recall (-3.3%), whereas LCA increases 0% recall precision by +5.9% (from 82.2% to 87.0%).

  8. Knowl 8 — Evaluation and Downweighting of Expansion Concepts on the WEST Legal Corpus

    data/table

    On the WEST legal collection (34 queries, 11,953 long documents averaging 1,970 words, baseline average precision of 53.8%), standard expansion techniques fail because the original queries are already high quality and average relevant documents per query is small (29 compared to 196 on TREC-3 and 133 on TREC-4). Standard local feedback causes an immediate loss in retrieval effectiveness across all feedback sizes (-7.8% at 5 docs to -34.7% at 100 docs).

    To prevent expansion concepts from degrading high-quality original queries, the weight of the auxiliary query Q′Q' is downweighted by 50% (w=1.0w = 1.0 instead of w=2.0w = 2.0).

    Recall Baseline Local Feedback (10 doc, dw0.5) LCA (100 passages, w=1.0w=1.0)
    0 88.0 81.9 (-7.0%) 92.1 (+4.7%)
    10 80.0 76.9 (-4.0%) 84.3 (+5.4%)
    20 77.5 71.4 (-7.8%) 78.5 (+1.3%)
    30 74.1 68.2 (-7.9%) 73.9 (-0.1%)
    40 62.9 60.8 (-3.3%) 61.8 (-1.7%)
    50 57.5 56.8 (-1.2%) 56.8 (-1.2%)
    60 49.7 50.1 (+0.8%) 50.7 (+2.2%)
    70 41.5 42.1 (+1.3%) 44.2 (+6.4%)
    80 32.7 33.1 (+1.1%) 36.4 (+11.2%)
    90 19.3 21.8 (+13.0%) 22.6 (+17.1%)
    100 8.6 9.3 (+7.8%) 10.0 (+15.3%)
    Average 53.8 52.0 (-3.3%) 55.6 (+3.3%)

    Even with 50% downweighting, local feedback fails to beat the baseline (-3.3% average precision). In contrast, LCA with w=1.0w=1.0 improves average precision to 55.6% (+3.3% at 100 passages, and +5.0% at 20 passages reaching 56.5%), boosting high-precision recall levels (88.0% to 92.1% at 0% recall).

  9. Knowl 9 — Robustness and Query-by-Query Degradation of LCA vs. Local Feedback

    empirical result

    A per-query analysis on the 49 TREC-4 queries demonstrates that Local Context Analysis (LCA) is far more robust and predictable than Rocchio-based local feedback:

    • Query Outcome Distribution: Out of 49 queries, local feedback improves 28 and hurts 21. In contrast, LCA improves 38 queries and hurts only 11.
    • Severe Degradation: Local feedback causes severe performance drops (>5% loss in average precision) on 5 queries, with the worst degradation reducing average precision from 24.8% to 4.3% (query 232). LCA produces a >5% drop on only 1 query.
    • Performance on Poor Baseline Queries: For the 9 queries with baseline average precision below 5%, local feedback degrades 8 queries and improves only 1. In contrast, LCA improves 5 of the 9 poorly performing queries and hurts 4.

    Local feedback relies heavily on high relevant document density in the top-ranked documents; when the initial retrieval contains few relevant items, local feedback incorporates non-relevant terms that compound retrieval failure. LCA mitigates this vulnerability by scoring co-occurrences against all query terms across passage units.

  10. Knowl 10 — Sensitivity of Expansion Effectiveness to Feedback Passage and Document Count

    empirical result

    The effectiveness of Local Context Analysis and local feedback exhibits distinct sensitivity profiles depending on the number of feedback units retrieved:

    • Local Context Analysis: On large corpora (TREC-3 and TREC-4), LCA is insensitive to the exact number of passages over a broad range: retrieval performance rises sharply from 10 to 50 passages, peaks around 100–200 passages (achieving 31.1% on TREC-4 and 39.3% on TREC-3), and remains stable across 30 to 300 passages before decaying slowly at 1,000–2,000 passages (27.9% on TREC-4 and 36.6% on TREC-3). On the smaller WEST corpus, LCA peaks earlier at 20 passages (56.5% with downweighting).
    • Local Feedback: Local feedback is sensitive to document count. On TREC-4, average precision falls monotonically from +14.0% improvement at 5 documents (28.7%) down to +3.5% at 100 documents (26.1%). On WEST, performance degradation intensifies from -7.8% at 5 documents to -34.7% at 100 documents (-2.2% to -25.7% with 50% downweighting).

    Thus, LCA provides a wide operational window (30–300 passages) without requiring fine-tuned document counts.

  11. Knowl 11 — Concept Vocabulary Divergence between Local Context Analysis and Local Feedback

    empirical result

    Despite both methods extracting expansion features from the top-ranked retrieved results of the same initial query, Local Context Analysis (LCA) and standard local feedback select largely disjoint sets of terms.

    On the 49 queries of TREC-4:

    • Standard local feedback adds an average of 58 unique terms per query (derived from 50 terms and 10 phrases).
    • LCA adds an average of 78 unique terms per query (derived from 70 top-scoring noun groups).
    • The average overlap between the two expansion vocabularies is only 17.6 terms per query.

    For example, on TREC-4 query 214 ("What are the different techniques used to create self induced hypnosis"), the overlap between LCA and local feedback concepts is 19 terms, yet both distinct expansions improve query performance over baseline. This shows that LCA and local feedback operate as fundamentally different expansion mechanisms rather than approximations of one another.

Coverage note — None was omitted; all key contributions—including the formal mathematical formulation of LCA, algorithmic procedure, comparison against global analysis (Phrasefinder) and local feedback across TREC-3, TREC-4, and WEST, parameter sensitivity analyses, robustness evaluations, and vocabulary overlap analysis—are fully covered.

References

  1. 1.Attar, R., & Fraenkel, A. S. (1977). Local Feedback in Full-Text Retrieval Systems. Journal of the Association for Computing Machinery, 24(3),397–417.
  2. 2.Buckley, C., Singhal, A., Mitra, M., & Salton, G. (1996). New Retrieval Approaches Using SMART : TREC 4. In Harman, D., editor, Proceedings of the TREC 4 Conference. National Institute of Standards and Technology Special Publication. to appear.
  3. 3.Caid, B., Gallant, S., Carleton, J., & Sudbeck, D. (1993). HNC Tipster Phase I Final Report. In Proceedings of Tipster Text Program (Phase I), pp. 69–92.
  4. 4.Callan, J., Croft, W. B., & Broglio, J. (1995). TREC and TIPSTER experiments with INQUERY. Information Processing and Management, pp. 327–343.
  5. 5.Callan, J. P. (1994). Passage-level evidence in document retrieval. In Proceedings of ACM SIGIR International Conference on Research and Development in Information Retrieval, pp. 302–310.
  6. 6.Croft, W. B., Cook, R., & Wilder, D. (1995). Providing Government Information on The Internet: Experiences with THOMAS. In Digital Libraries Conference DL '95, pp. 19–24.
  7. 7.Croft, W. B., & Harper, D. J. (1979). Using probabilistic models of document retrieval without relevance information. Journal of Documentation, 35,285–295.
  8. 8.Crouch, C. J., & Yang, B. (1992). Experiments in automatic statistical thesaurus construction. In Proceedings of ACM SIGIR International Conference on Research and Development in Information Retrieval, pp. 77–88.
  9. 9.Deerwester, S., Dumais, S., Furnas, G., Landauer, T., & Harshman, R. (1990). Indexing by latent semantic analysis. Journal of the American Society for Information Science, 41,391–407.
  10. 10.Harman, D. (1995). Overview of the Third Text REtrieval Conference (TREC-3). In Harman, D., editor, Proceedings of the Third Text REtrieval Conference (TREC-3), pp. 1–20. NIST Special Publication 500-225.
  11. 11.Harman, D., editor (1996). Proceedings of the TREC 4 Conference. National Institute of Standards and Technology Special Publication. to appear.
  12. 12.Jing, Y., & Croft, W. B. (1994). An association thesaurus for information retrieval. In Proceedings of RIAO 94, pp. 146–160.
  13. 13.Qiu, Y., & Frei, H. P. (1993). Concept based query expansion. In Proceedings of ACM SIGIR International Conference on Research and Development in Information Retrieval, pp. 160–169.
  14. 14.Schütze, H., & Pedersen, J. O. (1994). A cooccurrence-based thesaurus and two applications to information retrieval. In Proceedings of RIAO 94, pp. 266–274.
  15. 15.Sparck Jones, K. (1971). Automatic Keyword Classification for Information Retrieval. Butterworth, London.
  16. 16.Turtle, H. (1994). Natural language vs. Boolean query evaluation: A comparison of retrieval performance. In Proceedings of ACM SIGIR International Conference on Research and Development in Information Retrieval, pp. 212–220.
  17. 17.Voorhees, E. (1994). Query expansion using lexical-semantic relations. In Proceedings of ACM SIGIR International Conference on Research and Development in Information Retrieval, pp. 61–69.

Citation

MLA
Xu, J., and W. B. Croft. “Query Expansion Using Local and Global Document Analysis”. Proceedings of the 19th Annual International ACM SIGIR Conference on Research and Development in Information Retrieval - SIGIR '96, 1996, pp. 4–1, https://doi.org/10.1145/243199.243202.
APA
Xu, J., & Croft, W. B. (1996). Query expansion using local and global document analysis. Proceedings of the 19th Annual International ACM SIGIR Conference on Research and Development in Information Retrieval - SIGIR '96, 4–11. https://doi.org/10.1145/243199.243202
Chicago
Xu, J., and W. B. Croft. 1996. “Query Expansion Using Local and Global Document Analysis”. Proceedings of the 19th Annual International ACM SIGIR Conference on Research and Development in Information Retrieval - SIGIR '96, 4–11. https://doi.org/10.1145/243199.243202.
Harvard
Xu, J. and Croft, W.B. (1996) “Query expansion using local and global document analysis”, Proceedings of the 19th annual international ACM SIGIR conference on Research and development in information retrieval - SIGIR '96. ACM Press, pp. 4–11. Available at: https://doi.org/10.1145/243199.243202.
Vancouver
1. Xu J, Croft WB (1996) Query expansion using local and global document analysis. In: Proceedings of the 19th annual international ACM SIGIR conference on Research and development in information retrieval - SIGIR '96. ACM Press, pp 4–11

BibTeX

@inproceedings{Xu_1996, series={SIGIR ’96}, title={Query expansion using local and global document analysis}, url={http://dx.doi.org/10.1145/243199.243202}, DOI={10.1145/243199.243202}, booktitle={Proceedings of the 19th annual international ACM SIGIR conference on Research and development in information retrieval  - SIGIR ’96}, publisher={ACM Press}, author={Xu, Jinxi and Croft, W. Bruce}, year={1996}, pages={4–11}, collection={SIGIR ’96} }
Metadata:Crossref

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF