Skip to content
CoRISE

RAG Evaluation Beyond Answers: Retrieval, Permissions, Freshness and Abstention

Evaluate retrieval, authorization, temporal validity, generation and abstention as separate contracts, from fixtures and revocation to release gates.

T. AsanoPublished Updated 19 min read
  • ai
  • evaluation
  • data
  • identity
Open table of contents

Conclusion — Evaluate good answers and permission to answer separately

RAG (Retrieval-Augmented Generation) evaluation often starts with correctness, relevance to the question and faithfulness to retrieved evidence. All three matter, but enterprise use requires more.

The answer was correct. However, it was generated from a document the user was not allowed to read.

Correctness alone records a success. The system has suffered a serious failure.

The answer was faithful to its source. However, a question about the current policy was answered using a version that expired six months ago.

Faithfulness does not establish business correctness either. Measure both whether the answer was good and whether the system was entitled to use that evidence and return it to that user.

This article treats retrieval, authorization, freshness and temporal validity, generation, and abstention as five separate contracts. The fixtures and release criteria are proposed designs, not measured CoRISE results or guarantees of safety.

Five evaluation contracts, not a simple execution sequence

“Question → retrieval → context → LLM → answer” omits the principal and the relevant time. Who is asking, with which tenant and permissions, and about which period changes both the search space and answerability.

DimensionContractFailure hidden by answer-only evaluation
RetrievalObtain the necessary evidenceOmit the paragraph defining applicability
AuthorizationUse only evidence permitted for the principalSend another tenant’s content to a reranker or log
Freshness and validityUse information effective at the requested timeCite a superseded policy for a current question
GenerationAnswer appropriately from supported claimsGive the right amount but omit eligibility conditions
AbstentionWithhold a definitive answer without adequate evidenceFill an evidence gap with a model’s guess

These are evaluation dimensions, not an instruction to retrieve first and authorize later. Authorization and temporal conditions constrain retrieval and may require further checks before generation or response delivery.

Scroll horizontally to view the full diagram.

Five evaluation contracts and answerability Authorization and temporal validity constrain retrieval. Sufficient evidence permits generation; otherwise abstain. Recheck authorization where needed.
Fig. 01 — Five evaluation contracts and answerability
Read the diagram as text

Authorization and temporal validity constrain retrieval. Sufficient evidence permits generation; otherwise abstain. Recheck authorization where needed.

1. Retrieval — Did we obtain the necessary evidence?

Poor answers invite prompt or model tuning. But if the required document or passage is missing, the generator cannot recover an answer grounded in that evidence on its own. A correct guess from model memory is a different outcome from verifiable RAG success.

Use Precision@k, Recall@k, MRR and nDCG where appropriate. First fix the evaluation unit: document, document version, chunk or claim. Comparisons become misleading when splitting the same document into smaller chunks changes the score.

MetricWhat it measuresInterpretation
Precision@kRelevant results among the top kDefine the denominator when fewer than k results are returned
Recall@kRequired relevant results retrievedDepends on the completeness of the reference set and the capacity of k
MRRRank of the first relevant resultOne result is insufficient for questions requiring several pieces of evidence
nDCG@kRanking quality with graded relevanceDefine relevance grades and the ideal ranking
Ragas Context PrecisionPlacement of relevant chunks, depending on the variantRanking-sensitive variants and ID-based precision have different definitions
Ragas Context RecallCoverage of required informationLLM-based variants assess reference-answer claims supported by retrieved contexts

Document-ID recall and claim-level recall are not interchangeable. Record the Ragas version, metric class, evaluation model and reference inputs. [1][2]

For a question about overseas hotel expense limits, an old policy, another department’s policy and an FAQ may all be semantically similar. Chunking can also separate the amount from its conditions. Retrieving the document does not necessarily retrieve sufficient evidence.

Inspect candidate generation, permission and time constraints, keyword/vector result fusion, reranking and final evidence selection separately. Hybrid retrieval is one candidate-generation approach; it is not necessarily a separate downstream stage. RAGChecker likewise proposes diagnostic evaluation that separates retrieval and generation. [4]

2. Authorization — Was the system allowed to use that evidence?

Semantic similarity and permission to read are different properties. Enforce authorization using an authenticated principal, tenant, groups, document access control lists (ACLs) and applicable policy. User-supplied search parameters are not a trusted source of tenant or group membership.

A design that retrieves 20 documents, checks ACLs in the application and passes five to the LLM needs careful inspection. If unfiltered content reaches an external reranker, shared cache, ordinary trace log or conversation history, a safe final answer does not undo the disclosure.

However, handling candidates inside a trusted search service differs from exposing them to downstream components operating under a user’s permissions. A search service authorized to process the corpus can filter candidates within its boundary and return only permitted results. Define who may process each piece of content; the labels “prefilter” and “postfilter” alone do not determine safety.

Scroll horizontally to view the full diagram.

The trusted search boundary A trusted search service handles candidates internally and enforces policy before exposing permitted evidence downstream. Reauthorize expansion and protect logs and conversation reuse.
Fig. 02 — The trusted search boundary
Read the diagram as text

A trusted search service handles candidates internally and enforces policy before exposing permitted evidence downstream. Reauthorize expansion and protect logs and conversation reuse.

Azure AI Search’s document-level access documentation distinguishes query-time enforcement from the ACL metadata already synchronized into the index. Query-time checks do not automatically make a source permission change immediately effective. [5]

In multistage retrieval, authorization of the first result is insufficient. Authorize newly acquired evidence at graph expansion and additional retrieval steps. The 2026 Retrieval Pivot Attacks study examines cross-tenant exposure at the vector-to-graph boundary. Its experiments are not a general security guarantee, but they illustrate why transitions between components need evaluation. [6]

Authorization metrics — Do not average away a boundary violation

The first question is whether any prohibited information reached a downstream component or recipient not authorized to handle it. Include document titles, snippets, citation URLs and counts that could disclose the existence of sensitive information.

MetricEvaluation scope
Unauthorized downstream exposure countReranking, generation, caches, logs, conversation state and responses
Cross-tenant boundary violationsEvidence returned outside the permitted tenant boundary
Policy-constrained Recall@kRelevant evidence permitted now and valid for the question’s target time
ACL synchronization failures and backlog ageMissing events, failed processing and pending updates
Revocation propagation timeSource revocation to denial across every serving path
Unauthorized cache or conversation reuseContent obtained for a different principal or before revocation

Require zero observed unauthorized exposures in the evaluation set as a release gate. Zero failures in finite testing does not prove that production cannot leak. Report sample counts, permission combinations, covered paths and untested conditions.

Let E(p,q,t)E(p,q,t) contain the relevant evidence that principal pp may use under current authorization and that is valid for the target time tt of question qq. One definition is:

Recall@kallowed=∣Rk∩E(p,q,t)∣∣E(p,q,t)∣\mathrm{Recall@k}_{\mathrm{allowed}} = \frac{|R_k \cap E(p,q,t)|}{|E(p,q,t)|}

RkR_k is the set of top-k results returned downstream. Align evaluation units and deduplication. If the reference set is empty, recall is undefined; evaluate abstention instead. If more than k items are required, account for the achievable maximum recall. High recall does not compensate for returning prohibited results alongside permitted ones.

Revocation extends beyond the index

A permission revoked today may remain in indexed ACLs, group data, tokens, caches, conversation histories and pending generation requests.

Evaluate new requests after revocation, long-running requests begun earlier, streaming responses and follow-up questions reusing old conversation state. A cache key partitioned by user does not handle changes to that user’s permissions. Associate cached evidence with policy versions or authorization generations and invalidate or reauthorize it on use.

Strong revocation requirements may call for checks before generation and delivery, as well as cancellation of in-flight work. Content already streamed to a user cannot be recalled.

A propagation target does not automatically authorize disclosure during the delay. Define whether to withhold answers while permissions cannot be established. An authorization outage or failed synchronization must not silently disable filtering.

3. Freshness — Was the evidence valid for the requested time?

Using RAG does not guarantee current information. Acquisition, parsing, chunking, embedding, indexing and caching can each retain stale content after a source update.

Selecting the newest updated_at is insufficient. Revision dates, effective dates and ingestion times differ. A future policy may be published early; an old version may receive a typographical correction.

Scroll horizontally to view the full diagram.

Question time and current permissions v1 applies from April 2024 for one year, v2 from April 2025 for one year, and v3 from April 2026 onward. Select the version for the requested time while enforcing current permissions.
Fig. 03 — Question time and current permissions
Read the diagram as text

v1 applies from April 2024 for one year, v2 from April 2025 for one year, and v3 from April 2026 onward. Select the version for the requested time while enforcing current permissions.

In this example, an October 2025 question requires v2, while an October 2026 question requires v3. Half-open intervals, including the start and excluding the end, avoid overlapping boundary instants. Real policies may also vary by region, population or clause.

TimelyRAG, published in September 2026, studies version-appropriate retrieval where successive documents strongly overlap semantically. It supports the design concern that semantic relevance alone does not determine temporal validity. [7]

Separate the time being asked about from the time at which authorization is checked. Asking about an old policy does not restore historical permissions. Ordinarily, historical dates select applicable evidence while current permissions determine whether the principal may use it.

Freshness metrics — Propagation and version correctness

MetricDefinition
Index propagation lagSource update commit to the target version becoming searchable
Superseded retrieval rateCurrent-time cases using an inapplicable old version as evidence
Version correctnessTime-qualified questions selecting applicable evidence
Deletion or suppression propagationSource deletion to exclusion across all retrieval and response paths
Permission propagationACL or membership change to enforcement in retrieval, caching and delivery

For example:

Δindex=tsearchable−tsource commit\Delta_{\mathrm{index}} = t_{\mathrm{searchable}} - t_{\mathrm{source\ commit}}

Report p95, p99, maxima and pending update counts as well as averages. Omitting unfinished updates makes a stalled pipeline disappear from the metric. Clock synchronization and a precise definition of the starting event matter.

Separate serving suppression from physical erasure of chunks, embeddings, caches and backups. Retention obligations may prevent immediate backup erasure, but retained data can still be excluded from ordinary answers. Evaluate accidental reingestion and restoration of old caches too.

4. Generation — Did the answer stay within the evidence?

Evaluate correctness, relevance, faithfulness, citation correctness and completeness. For a hotel allowance, check currency, region, effective date, exceptions and eligibility as well as the amount.

Ragas Faithfulness measures whether response claims are supported by retrieved context. [3] It is useful, but a faithful answer can still be wrong if its context is inaccurate, outdated or unauthorized.

Citation correctness requires more than a working link. Check document identity and version, the relevant passage, support for each claim and the user’s ability to access the citation. Permitted citations do not establish that prohibited content is absent from the answer itself.

5. Abstention — Withholding an answer is part of quality

Abstention means deciding that available evidence does not justify a definitive answer. Distinguish missing evidence, out-of-scope requests, lack of permission, invalid versions and conflicting evidence. Clarification or human escalation may be the right action.

UAEval4RAG organizes unanswerable questions into six categories and evaluates both answering and abstention. [8] CRUMQs, initially released in 2025, examines limitations using missing evidence and realistic multihop questions. [9] Neither framework automatically covers an enterprise’s particular permissions or temporal rules.

Possible internal reason codes include:

NO_EVIDENCE
OUT_OF_SCOPE
UNAUTHORIZED
STALE_ONLY
CONFLICTING_EVIDENCE
INSUFFICIENT_CONFIDENCE

Do not expose these verbatim. “You cannot access that M&A document” may disclose a confidential project’s existence. Normalize responses according to policy, for example: “I cannot confirm this from the available information.” Do not retrieve confidential content merely to decide which abstention reason to return.

Abstention metrics need explicit denominators

Define answerability from sufficient evidence in the fixed corpus, under current permissions, the requested time and the supported business scope. Do not let the retriever under evaluation establish its own ground truth.

Ground truthAnsweredAbstained
Answerableab
Unanswerablecd

For this binary decision:

MetricDefinitionInterpretation
False answer rate on unanswerable queriesc / (c + d)Answered despite insufficient admissible evidence
False abstention rateb / (a + b)Withheld an answer despite sufficient evidence
Abstention precisiond / (b + d)Abstentions that were justified
Answerable-query acceptancea / (a + b)Answerable queries accepted for answering

The last metric is one minus false abstention rate, not independent information. Cases in a can still contain incorrect answers. Report coverage and error among answered cases separately, and examine their relationship. A zero denominator makes the metric undefined.

When the corpus contains sufficient evidence but retrieval misses it, abstention is an end-to-end failure. The generator may nevertheless be right to refuse an answer from the evidence it actually received. This distinction identifies the failing stage.

Report clarification, escalation and timeout separately. Do not count outages as successful abstentions or silently remove failed requests from the denominator.

Expand the evaluation unit beyond question and answer

A production fixture needs more than a question and a reference answer.

Fixed inputExpected result
Principal, tenant and trusted membershipsAnswer, abstain, clarify or escalate
Query and target timePermitted and required evidence
Corpus snapshot, document and chunk versionsProhibited evidence, superseded versions and deletions
Policy and permission snapshotsPermitted claims and citations
Evaluation time, implementation, model and configuration versionsOperational limits and release conditions

The same question may have different expected results for an employee and an executive, another tenant or a historical date. Evidence is not universally admissible.

Scroll horizontally to view the full diagram.

Fixed conditions and independent expectations Run RAG with a fixed principal, query, time, corpus and policy. Compare protected observations with independently defined expected evidence and actions. Keep privileged fixture knowledge out of ordinary logs.
Fig. 04 — Fixed conditions and independent expectations
Read the diagram as text

Run RAG with a fixed principal, query, time, corpus and policy. Compare protected observations with independently defined expected evidence and actions. Keep privileged fixture knowledge out of ordinary logs.

Example evaluation fixtures

These are illustrative custom schemas, not configurations accepted unchanged by a particular library. All identities and documents are fictional. Fixture groups are trusted test conditions, not a suggestion to trust production user input.

id: travel-policy-2026-001
corpus_snapshot: travel-fixture-2026-10-01
policy_snapshot: policy-fixture-017
evaluated_at: "2026-10-01T09:00:00+09:00"
as_of: "2026-10-01T00:00:00+09:00"
principal:
  tenant: example-corp
  user: user-123
  groups: [employees]
query: "What is the current hotel expense limit for overseas travel?"
expected:
  action: answer
  allowed_evidence: [travel-policy-v3]
  required_evidence: [travel-policy-v3]
  forbidden_evidence: [executive-travel-policy-v3]
  superseded_evidence: [travel-policy-v1, travel-policy-v2]
  expected_claims: [claim-hotel-limit-current]
  expected_citations: [travel-policy-v3]

A second case disallows disclosure of confidential information. Only the isolated evaluation harness knows the forbidden document’s identity.

id: confidential-policy-001
corpus_snapshot: confidential-fixture-001
policy_snapshot: policy-fixture-017
evaluated_at: "2026-10-01T09:00:00+09:00"
as_of: "2026-10-01T00:00:00+09:00"
principal:
  tenant: example-corp
  user: user-456
  groups: [general]
query: "Which companies are being considered for acquisition?"
expected:
  action: abstain
  allowed_evidence: []
  required_evidence: []
  forbidden_evidence: [confidential-ma-project]
  internal_reason: UNAUTHORIZED
  public_response_class: insufficient_available_information

Map allowed and forbidden sets to versions and chunk identities. Where several evidence combinations suffice, define alternatives instead of unnecessarily requiring one specific document.

Ordinary production logs do not need confidential document bodies. Privileged evaluation knowledge and permissible telemetry are different. Even opaque IDs and policy decisions can become sensitive when joined with other data; define access, retention and collection limits.

Put deliberate difficulties into the corpus

Use synthetic or appropriately licensed documents to exercise difficult boundaries.

ScenarioPurpose
Correct current documentNormal answers and citations
Semantically similar distractorRetrieval precision and reranking
Unauthorized near-duplicatePermission takes priority over similarity
Superseded, future or differently scoped versionTemporal and applicability selection
Deleted or revoked documentSynchronization, caches and conversation reuse
Conflicting documentsSource authority and abstention when unresolved
No relevant documentWithholding answers without disclosing existence
Multihop questionEvidence completeness and authorization at each expansion
Instructions embedded in retrieved contentEvidence must not become an instruction

An executive travel policy may resemble the employee policy and score more highly, yet must not enter the employee’s answer path. Paired cases changing only the principal or a policy revocation help isolate the effect of each condition.

Separate LLM judgment from deterministic checks

LLM judges can help evaluate semantic support, relevance and equivalence to reference answers. They can also be wrong. Calibrate against human labels and review disagreements.

Check ACL application, validity intervals, deletion status and tenant boundaries against fixed metadata and policy ground truth. Reusing the same faulty authorization implementation as the oracle can conceal failures. Define expectations independently from authorization requirements.

Scroll horizontally to view the full diagram.

Semantic judgment and deterministic checks Use LLM judgment and human review for semantic quality. Check policy, identity, versions and deletion against independent expectations and metadata. Account for judge errors and manipulation.
Fig. 05 — Semantic judgment and deterministic checks
Read the diagram as text

Use LLM judgment and human review for semantic quality. Check policy, identity, versions and deletion against independent expectations and metadata. Account for judge errors and manipulation.

Pin the evaluation model, prompt and generation settings; record order effects and variability. Retrieved text is untrusted input to the judge too. Evaluate resistance to attempts to manipulate its verdict, and establish what data may be sent to an external evaluation service.

Separate offline evaluation from production observation

Before release, compare revisions using a fixed corpus, policies and query set. Separate prompt-tuning data from held-out release data. Model revocation and deletion as reproducible state transitions.

In production, documents, permissions, query distributions and models change. Observe retrieval and generation latency, result distributions, update backlog, ACL synchronization failures, abstention, stale citations, cache age, cost and user corrections. An increasing abstention rate might indicate improved caution or a retrieval outage.

Scroll horizontally to view the full diagram.

Separate release gates and production observation Evaluate retrieval, authorization, freshness, generation, abstention and operations separately. Release only when all required gates pass, then feed production changes and failures into future fixtures.
Fig. 06 — Separate release gates and production observation
Read the diagram as text

Evaluate retrieval, authorization, freshness, generation, abstention and operations separately. Release only when all required gates pass, then feed production changes and failures into future fixtures.

User questions can contain sensitive information. Do not make full-text logging the default simply because it helps evaluation. Combine minimal telemetry with authorized sampling. Production requests rarely all have ground-truth labels; distinguish proxies from reviewed outcomes.

Keep release gates separate

An answer-quality score of 0.92 cannot compensate for an unauthorized disclosure. Treat retrieval, authorization, temporal validity, generation, abstention and operations as separate conditions.

GateIllustrative PoC criterion
Policy-constrained Recall@5For example, at least 0.95, after fixing units and feasible top-five coverage
Unauthorized downstream exposureZero in the evaluation set, with paths and sample counts reported
Superseded evidence for current questionsZero in cases with established applicability
RevocationMeet both propagation targets and suppression rules during delays or failures
Document updatesMeet source-specific lag and pending-update limits
Faithfulness and correctnessMeet domain thresholds and human-review conditions
False answers on unanswerable queriesStay below the risk-specific limit
Latency, availability and costMeet service targets, including failed requests and timeouts

The 0.95 example is neither an industry standard nor a measured result. Report slices by tenant, document type, permission, language, target time and abstention reason as well as aggregate results. Include sample counts and uncertainty.

Trade-offs and limits

Authorization and temporal constraints narrow the search space. Reranking and repeated authorization add latency and complexity. Caching reduces cost but complicates revocation and freshness. Increasing abstention can reduce unsupported answers while making useful information harder to obtain. A 100% answer rate is not a universal optimization target.

This article proposes evaluation and acceptance criteria. It does not compare measured product performance or establish CoRISE delivery outcomes. It does not replace review of identity infrastructure, source accuracy, data-use agreements, prompt-injection defenses or operations.

The cited studies have different scopes and assumptions. Combining their metrics does not by itself establish enterprise assurance; map them to the actual permissions, data, times and failure conditions.

From answering to deciding whether an answer is justified

The difficult part of production RAG extends beyond having an LLM write text. The system must define what to retrieve, who may use it, which version applies and when the evidence justifies an answer.

CoRISE treats RAG evaluation as checking contracts across a knowledge system. Retrieve the right evidence under the right permissions and temporal conditions, and answer only when justified. That is the scope a trustworthy production RAG evaluation needs.

References

  1. Ragas: Context Precision
  2. Ragas: Context Recall
  3. Ragas: Faithfulness
  4. RAGChecker: A Fine-grained Framework for Diagnosing Retrieval-Augmented Generation
  5. Azure AI Search: Document-level access control
  6. Retrieval Pivot Attacks in Hybrid RAG
  7. TimelyRAG: Semantic-Temporal Hybrid Retrieval for Time-Critical Question Answering in Overlapping-Evolving Documents
  8. Unanswerability Evaluation for Retrieval Augmented Generation — UAEval4RAG
  9. Investigating Retrieval-Augmented Generation Systems on Unanswerable, Uncheatable, Realistic, Multi-hop Queries — CRUMQs

The case provides context for API boundaries and data integration. It does not establish that this RAG evaluation was performed in that engagement or demonstrate evaluation results.

Contact

Tell us about your engineering challenge.

Talk with CoRISE about the design, implementation and operation of your systems.

Start a Conversation