Skip to content
CoRISE

Comparing Vector Retrieval for Long Enterprise Documents

Compare long-document retrieval under explicit chunking, evaluation and query conditions.

2 min read
  • Data
  • Evaluation
  • AI
Open table of contents

Conclusion

Compare long-document retrieval by whether it returns the evidence needed for the same questions. First hold the corpus and chunking constant and vary the retrieval method. Change chunking and reranking separately afterward so their effects remain distinguishable.

Candidates and evaluation units

Consider keyword retrieval, vector retrieval and a combination of the two. Include exact identifiers, paraphrases, questions requiring evidence from several passages and questions with no answer in the corpus.

Label source sections or passages as evidence, rather than only document names. Retrieving an unrelated passage from the right document is not necessarily a success. Group overlapping chunks by source and location so repeated retrieval of the same evidence does not inflate the score.

A repeatable comparison

  1. Separate development and final evaluation documents and questions. Split by source-document family so near-duplicate questions and revisions of the same source do not leak across the boundary.
  2. Record source revisions, chunk length and overlap, embedding model, retrieval settings and evaluation-script revision.
  3. Run the same queries through each method and retain candidates, scores and retrieval duration. Keep cache conditions and concurrency comparable.
  4. Calculate measures such as Recall@k for answerable questions and inspect failures by question type. The denominator for Recall@k is the total number of labeled relevant evidence units. Assess unanswerable questions separately, including how often the system incorrectly treats a result as supporting evidence.
  5. Also compare with an equal total token budget for material sent to generation. Equal candidate counts do not provide equal context when chunk sizes differ.
  6. Select settings on the development set and run the fixed comparison on the final evaluation set. Retain failed cases and configuration. If those results lead to further tuning, treat that work as development again.

Trade-offs

Larger chunks can preserve surrounding context while adding irrelevant material. Smaller chunks support localized evidence but may separate a statement from its conditions. Reranking adds a decision stage and processing cost, so compare quality and waiting time together.

Limitations

Retrieval measures alone do not establish answer correctness or access protection. Separately enforce the eligible document set and assess whether generated answers follow the evidence.

Sources

These cases provide attributed design context. They do not establish that the proposed experiments or configurations were delivered in those engagements.

Contact

Tell us about your engineering challenge.

Talk with CoRISE about the design, implementation and operation of your systems.

Start a Conversation