Open table of contents
Conclusion — Evaluate good answers and permission to answer separately
RAG (Retrieval-Augmented Generation) evaluation often starts with correctness, relevance to the question and faithfulness to retrieved evidence. All three matter, but enterprise use requires more.
The answer was correct. However, it was generated from a document the user was not allowed to read.
Correctness alone records a success. The system has suffered a serious failure.
The answer was faithful to its source. However, a question about the current policy was answered using a version that expired six months ago.
Faithfulness does not establish business correctness either. Measure both whether the answer was good and whether the system was entitled to use that evidence and return it to that user.
This article treats retrieval, authorization, freshness and temporal validity, generation, and abstention as five separate contracts. The fixtures and release criteria are proposed designs, not measured CoRISE results or guarantees of safety.
Five evaluation contracts, not a simple execution sequence
“Question → retrieval → context → LLM → answer” omits the principal and the relevant time. Who is asking, with which tenant and permissions, and about which period changes both the search space and answerability.
| Dimension | Contract | Failure hidden by answer-only evaluation |
|---|---|---|
| Retrieval | Obtain the necessary evidence | Omit the paragraph defining applicability |
| Authorization | Use only evidence permitted for the principal | Send another tenant’s content to a reranker or log |
| Freshness and validity | Use information effective at the requested time | Cite a superseded policy for a current question |
| Generation | Answer appropriately from supported claims | Give the right amount but omit eligibility conditions |
| Abstention | Withhold a definitive answer without adequate evidence | Fill an evidence gap with a model’s guess |
These are evaluation dimensions, not an instruction to retrieve first and authorize later. Authorization and temporal conditions constrain retrieval and may require further checks before generation or response delivery.
Scroll horizontally to view the full diagram.
Read the diagram as text
Authorization and temporal validity constrain retrieval. Sufficient evidence permits generation; otherwise abstain. Recheck authorization where needed.
1. Retrieval — Did we obtain the necessary evidence?
Poor answers invite prompt or model tuning. But if the required document or passage is missing, the generator cannot recover an answer grounded in that evidence on its own. A correct guess from model memory is a different outcome from verifiable RAG success.
Use Precision@k, Recall@k, MRR and nDCG where appropriate. First fix the evaluation unit: document, document version, chunk or claim. Comparisons become misleading when splitting the same document into smaller chunks changes the score.
| Metric | What it measures | Interpretation |
|---|---|---|
| Precision@k | Relevant results among the top k | Define the denominator when fewer than k results are returned |
| Recall@k | Required relevant results retrieved | Depends on the completeness of the reference set and the capacity of k |
| MRR | Rank of the first relevant result | One result is insufficient for questions requiring several pieces of evidence |
| nDCG@k | Ranking quality with graded relevance | Define relevance grades and the ideal ranking |
| Ragas Context Precision | Placement of relevant chunks, depending on the variant | Ranking-sensitive variants and ID-based precision have different definitions |
| Ragas Context Recall | Coverage of required information | LLM-based variants assess reference-answer claims supported by retrieved contexts |
Document-ID recall and claim-level recall are not interchangeable. Record the Ragas version, metric class, evaluation model and reference inputs. [1][2]
For a question about overseas hotel expense limits, an old policy, another department’s policy and an FAQ may all be semantically similar. Chunking can also separate the amount from its conditions. Retrieving the document does not necessarily retrieve sufficient evidence.
Inspect candidate generation, permission and time constraints, keyword/vector result fusion, reranking and final evidence selection separately. Hybrid retrieval is one candidate-generation approach; it is not necessarily a separate downstream stage. RAGChecker likewise proposes diagnostic evaluation that separates retrieval and generation. [4]
2. Authorization — Was the system allowed to use that evidence?
Semantic similarity and permission to read are different properties. Enforce authorization using an authenticated principal, tenant, groups, document access control lists (ACLs) and applicable policy. User-supplied search parameters are not a trusted source of tenant or group membership.
A design that retrieves 20 documents, checks ACLs in the application and passes five to the LLM needs careful inspection. If unfiltered content reaches an external reranker, shared cache, ordinary trace log or conversation history, a safe final answer does not undo the disclosure.
However, handling candidates inside a trusted search service differs from exposing them to downstream components operating under a user’s permissions. A search service authorized to process the corpus can filter candidates within its boundary and return only permitted results. Define who may process each piece of content; the labels “prefilter” and “postfilter” alone do not determine safety.
Scroll horizontally to view the full diagram.
Read the diagram as text
A trusted search service handles candidates internally and enforces policy before exposing permitted evidence downstream. Reauthorize expansion and protect logs and conversation reuse.
Azure AI Search’s document-level access documentation distinguishes query-time enforcement from the ACL metadata already synchronized into the index. Query-time checks do not automatically make a source permission change immediately effective. [5]
In multistage retrieval, authorization of the first result is insufficient. Authorize newly acquired evidence at graph expansion and additional retrieval steps. The 2026 Retrieval Pivot Attacks study examines cross-tenant exposure at the vector-to-graph boundary. Its experiments are not a general security guarantee, but they illustrate why transitions between components need evaluation. [6]
Authorization metrics — Do not average away a boundary violation
The first question is whether any prohibited information reached a downstream component or recipient not authorized to handle it. Include document titles, snippets, citation URLs and counts that could disclose the existence of sensitive information.
| Metric | Evaluation scope |
|---|---|
| Unauthorized downstream exposure count | Reranking, generation, caches, logs, conversation state and responses |
| Cross-tenant boundary violations | Evidence returned outside the permitted tenant boundary |
| Policy-constrained Recall@k | Relevant evidence permitted now and valid for the question’s target time |
| ACL synchronization failures and backlog age | Missing events, failed processing and pending updates |
| Revocation propagation time | Source revocation to denial across every serving path |
| Unauthorized cache or conversation reuse | Content obtained for a different principal or before revocation |
Require zero observed unauthorized exposures in the evaluation set as a release gate. Zero failures in finite testing does not prove that production cannot leak. Report sample counts, permission combinations, covered paths and untested conditions.
Let contain the relevant evidence that principal may use under current authorization and that is valid for the target time of question . One definition is:
is the set of top-k results returned downstream. Align evaluation units and deduplication. If the reference set is empty, recall is undefined; evaluate abstention instead. If more than k items are required, account for the achievable maximum recall. High recall does not compensate for returning prohibited results alongside permitted ones.
Revocation extends beyond the index
A permission revoked today may remain in indexed ACLs, group data, tokens, caches, conversation histories and pending generation requests.
Evaluate new requests after revocation, long-running requests begun earlier, streaming responses and follow-up questions reusing old conversation state. A cache key partitioned by user does not handle changes to that user’s permissions. Associate cached evidence with policy versions or authorization generations and invalidate or reauthorize it on use.
Strong revocation requirements may call for checks before generation and delivery, as well as cancellation of in-flight work. Content already streamed to a user cannot be recalled.
A propagation target does not automatically authorize disclosure during the delay. Define whether to withhold answers while permissions cannot be established. An authorization outage or failed synchronization must not silently disable filtering.
3. Freshness — Was the evidence valid for the requested time?
Using RAG does not guarantee current information. Acquisition, parsing, chunking, embedding, indexing and caching can each retain stale content after a source update.
Selecting the newest updated_at is insufficient. Revision dates, effective dates and ingestion times differ. A future policy may be published early; an old version may receive a typographical correction.
Scroll horizontally to view the full diagram.
Read the diagram as text
v1 applies from April 2024 for one year, v2 from April 2025 for one year, and v3 from April 2026 onward. Select the version for the requested time while enforcing current permissions.
In this example, an October 2025 question requires v2, while an October 2026 question requires v3. Half-open intervals, including the start and excluding the end, avoid overlapping boundary instants. Real policies may also vary by region, population or clause.
TimelyRAG, published in September 2026, studies version-appropriate retrieval where successive documents strongly overlap semantically. It supports the design concern that semantic relevance alone does not determine temporal validity. [7]
Separate the time being asked about from the time at which authorization is checked. Asking about an old policy does not restore historical permissions. Ordinarily, historical dates select applicable evidence while current permissions determine whether the principal may use it.
Freshness metrics — Propagation and version correctness
| Metric | Definition |
|---|---|
| Index propagation lag | Source update commit to the target version becoming searchable |
| Superseded retrieval rate | Current-time cases using an inapplicable old version as evidence |
| Version correctness | Time-qualified questions selecting applicable evidence |
| Deletion or suppression propagation | Source deletion to exclusion across all retrieval and response paths |
| Permission propagation | ACL or membership change to enforcement in retrieval, caching and delivery |
For example:
Report p95, p99, maxima and pending update counts as well as averages. Omitting unfinished updates makes a stalled pipeline disappear from the metric. Clock synchronization and a precise definition of the starting event matter.
Separate serving suppression from physical erasure of chunks, embeddings, caches and backups. Retention obligations may prevent immediate backup erasure, but retained data can still be excluded from ordinary answers. Evaluate accidental reingestion and restoration of old caches too.
4. Generation — Did the answer stay within the evidence?
Evaluate correctness, relevance, faithfulness, citation correctness and completeness. For a hotel allowance, check currency, region, effective date, exceptions and eligibility as well as the amount.
Ragas Faithfulness measures whether response claims are supported by retrieved context. [3] It is useful, but a faithful answer can still be wrong if its context is inaccurate, outdated or unauthorized.
Citation correctness requires more than a working link. Check document identity and version, the relevant passage, support for each claim and the user’s ability to access the citation. Permitted citations do not establish that prohibited content is absent from the answer itself.
5. Abstention — Withholding an answer is part of quality
Abstention means deciding that available evidence does not justify a definitive answer. Distinguish missing evidence, out-of-scope requests, lack of permission, invalid versions and conflicting evidence. Clarification or human escalation may be the right action.
UAEval4RAG organizes unanswerable questions into six categories and evaluates both answering and abstention. [8] CRUMQs, initially released in 2025, examines limitations using missing evidence and realistic multihop questions. [9] Neither framework automatically covers an enterprise’s particular permissions or temporal rules.
Possible internal reason codes include:
NO_EVIDENCE
OUT_OF_SCOPE
UNAUTHORIZED
STALE_ONLY
CONFLICTING_EVIDENCE
INSUFFICIENT_CONFIDENCE
Do not expose these verbatim. “You cannot access that M&A document” may disclose a confidential project’s existence. Normalize responses according to policy, for example: “I cannot confirm this from the available information.” Do not retrieve confidential content merely to decide which abstention reason to return.
Abstention metrics need explicit denominators
Define answerability from sufficient evidence in the fixed corpus, under current permissions, the requested time and the supported business scope. Do not let the retriever under evaluation establish its own ground truth.
| Ground truth | Answered | Abstained |
|---|---|---|
| Answerable | a | b |
| Unanswerable | c | d |
For this binary decision:
| Metric | Definition | Interpretation |
|---|---|---|
| False answer rate on unanswerable queries | c / (c + d) | Answered despite insufficient admissible evidence |
| False abstention rate | b / (a + b) | Withheld an answer despite sufficient evidence |
| Abstention precision | d / (b + d) | Abstentions that were justified |
| Answerable-query acceptance | a / (a + b) | Answerable queries accepted for answering |
The last metric is one minus false abstention rate, not independent information. Cases in a can still contain incorrect answers. Report coverage and error among answered cases separately, and examine their relationship. A zero denominator makes the metric undefined.
When the corpus contains sufficient evidence but retrieval misses it, abstention is an end-to-end failure. The generator may nevertheless be right to refuse an answer from the evidence it actually received. This distinction identifies the failing stage.
Report clarification, escalation and timeout separately. Do not count outages as successful abstentions or silently remove failed requests from the denominator.
Expand the evaluation unit beyond question and answer
A production fixture needs more than a question and a reference answer.
| Fixed input | Expected result |
|---|---|
| Principal, tenant and trusted memberships | Answer, abstain, clarify or escalate |
| Query and target time | Permitted and required evidence |
| Corpus snapshot, document and chunk versions | Prohibited evidence, superseded versions and deletions |
| Policy and permission snapshots | Permitted claims and citations |
| Evaluation time, implementation, model and configuration versions | Operational limits and release conditions |
The same question may have different expected results for an employee and an executive, another tenant or a historical date. Evidence is not universally admissible.
Scroll horizontally to view the full diagram.
Read the diagram as text
Run RAG with a fixed principal, query, time, corpus and policy. Compare protected observations with independently defined expected evidence and actions. Keep privileged fixture knowledge out of ordinary logs.
Example evaluation fixtures
These are illustrative custom schemas, not configurations accepted unchanged by a particular library. All identities and documents are fictional. Fixture groups are trusted test conditions, not a suggestion to trust production user input.
id: travel-policy-2026-001
corpus_snapshot: travel-fixture-2026-10-01
policy_snapshot: policy-fixture-017
evaluated_at: "2026-10-01T09:00:00+09:00"
as_of: "2026-10-01T00:00:00+09:00"
principal:
tenant: example-corp
user: user-123
groups: [employees]
query: "What is the current hotel expense limit for overseas travel?"
expected:
action: answer
allowed_evidence: [travel-policy-v3]
required_evidence: [travel-policy-v3]
forbidden_evidence: [executive-travel-policy-v3]
superseded_evidence: [travel-policy-v1, travel-policy-v2]
expected_claims: [claim-hotel-limit-current]
expected_citations: [travel-policy-v3]
A second case disallows disclosure of confidential information. Only the isolated evaluation harness knows the forbidden document’s identity.
id: confidential-policy-001
corpus_snapshot: confidential-fixture-001
policy_snapshot: policy-fixture-017
evaluated_at: "2026-10-01T09:00:00+09:00"
as_of: "2026-10-01T00:00:00+09:00"
principal:
tenant: example-corp
user: user-456
groups: [general]
query: "Which companies are being considered for acquisition?"
expected:
action: abstain
allowed_evidence: []
required_evidence: []
forbidden_evidence: [confidential-ma-project]
internal_reason: UNAUTHORIZED
public_response_class: insufficient_available_information
Map allowed and forbidden sets to versions and chunk identities. Where several evidence combinations suffice, define alternatives instead of unnecessarily requiring one specific document.
Ordinary production logs do not need confidential document bodies. Privileged evaluation knowledge and permissible telemetry are different. Even opaque IDs and policy decisions can become sensitive when joined with other data; define access, retention and collection limits.
Put deliberate difficulties into the corpus
Use synthetic or appropriately licensed documents to exercise difficult boundaries.
| Scenario | Purpose |
|---|---|
| Correct current document | Normal answers and citations |
| Semantically similar distractor | Retrieval precision and reranking |
| Unauthorized near-duplicate | Permission takes priority over similarity |
| Superseded, future or differently scoped version | Temporal and applicability selection |
| Deleted or revoked document | Synchronization, caches and conversation reuse |
| Conflicting documents | Source authority and abstention when unresolved |
| No relevant document | Withholding answers without disclosing existence |
| Multihop question | Evidence completeness and authorization at each expansion |
| Instructions embedded in retrieved content | Evidence must not become an instruction |
An executive travel policy may resemble the employee policy and score more highly, yet must not enter the employee’s answer path. Paired cases changing only the principal or a policy revocation help isolate the effect of each condition.
Separate LLM judgment from deterministic checks
LLM judges can help evaluate semantic support, relevance and equivalence to reference answers. They can also be wrong. Calibrate against human labels and review disagreements.
Check ACL application, validity intervals, deletion status and tenant boundaries against fixed metadata and policy ground truth. Reusing the same faulty authorization implementation as the oracle can conceal failures. Define expectations independently from authorization requirements.
Scroll horizontally to view the full diagram.
Read the diagram as text
Use LLM judgment and human review for semantic quality. Check policy, identity, versions and deletion against independent expectations and metadata. Account for judge errors and manipulation.
Pin the evaluation model, prompt and generation settings; record order effects and variability. Retrieved text is untrusted input to the judge too. Evaluate resistance to attempts to manipulate its verdict, and establish what data may be sent to an external evaluation service.
Separate offline evaluation from production observation
Before release, compare revisions using a fixed corpus, policies and query set. Separate prompt-tuning data from held-out release data. Model revocation and deletion as reproducible state transitions.
In production, documents, permissions, query distributions and models change. Observe retrieval and generation latency, result distributions, update backlog, ACL synchronization failures, abstention, stale citations, cache age, cost and user corrections. An increasing abstention rate might indicate improved caution or a retrieval outage.
Scroll horizontally to view the full diagram.
Read the diagram as text
Evaluate retrieval, authorization, freshness, generation, abstention and operations separately. Release only when all required gates pass, then feed production changes and failures into future fixtures.
User questions can contain sensitive information. Do not make full-text logging the default simply because it helps evaluation. Combine minimal telemetry with authorized sampling. Production requests rarely all have ground-truth labels; distinguish proxies from reviewed outcomes.
Keep release gates separate
An answer-quality score of 0.92 cannot compensate for an unauthorized disclosure. Treat retrieval, authorization, temporal validity, generation, abstention and operations as separate conditions.
| Gate | Illustrative PoC criterion |
|---|---|
| Policy-constrained Recall@5 | For example, at least 0.95, after fixing units and feasible top-five coverage |
| Unauthorized downstream exposure | Zero in the evaluation set, with paths and sample counts reported |
| Superseded evidence for current questions | Zero in cases with established applicability |
| Revocation | Meet both propagation targets and suppression rules during delays or failures |
| Document updates | Meet source-specific lag and pending-update limits |
| Faithfulness and correctness | Meet domain thresholds and human-review conditions |
| False answers on unanswerable queries | Stay below the risk-specific limit |
| Latency, availability and cost | Meet service targets, including failed requests and timeouts |
The 0.95 example is neither an industry standard nor a measured result. Report slices by tenant, document type, permission, language, target time and abstention reason as well as aggregate results. Include sample counts and uncertainty.
Trade-offs and limits
Authorization and temporal constraints narrow the search space. Reranking and repeated authorization add latency and complexity. Caching reduces cost but complicates revocation and freshness. Increasing abstention can reduce unsupported answers while making useful information harder to obtain. A 100% answer rate is not a universal optimization target.
This article proposes evaluation and acceptance criteria. It does not compare measured product performance or establish CoRISE delivery outcomes. It does not replace review of identity infrastructure, source accuracy, data-use agreements, prompt-injection defenses or operations.
The cited studies have different scopes and assumptions. Combining their metrics does not by itself establish enterprise assurance; map them to the actual permissions, data, times and failure conditions.
From answering to deciding whether an answer is justified
The difficult part of production RAG extends beyond having an LLM write text. The system must define what to retrieve, who may use it, which version applies and when the evidence justifies an answer.
CoRISE treats RAG evaluation as checking contracts across a knowledge system. Retrieve the right evidence under the right permissions and temporal conditions, and answer only when justified. That is the scope a trustworthy production RAG evaluation needs.
References
- Ragas: Context Precision
- Ragas: Context Recall
- Ragas: Faithfulness
- RAGChecker: A Fine-grained Framework for Diagnosing Retrieval-Augmented Generation
- Azure AI Search: Document-level access control
- Retrieval Pivot Attacks in Hybrid RAG
- TimelyRAG: Semantic-Temporal Hybrid Retrieval for Time-Critical Question Answering in Overlapping-Evolving Documents
- Unanswerability Evaluation for Retrieval Augmented Generation — UAEval4RAG
- Investigating Retrieval-Augmented Generation Systems on Unanswerable, Uncheatable, Realistic, Multi-hop Queries — CRUMQs
Related case context
The case provides context for API boundaries and data integration. It does not establish that this RAG evaluation was performed in that engagement or demonstrate evaluation results.