Open table of contents
Conclusion
RAG and fine-tuning can address overlapping goals, but change different parts of a system. Separate a need to refresh reference information from a need to change response format or behavior. Knowledge encoded in a model does not automatically provide source attribution or access enforcement.
Four comparison conditions
Start with the same workflow and base model. Check whether that model supports the proposed training method before defining the comparison.
| Condition | What changes | Main evaluation questions |
|---|---|---|
| Baseline | Prompt only | Quality achievable without additional components |
| RAG | Retrieval and reference material | Freshness, attribution and retrieval failures |
| Fine-tuning | Training examples | Format compliance, consistent decisions and generalization |
| Combined | Retrieval and additional training | Combined effects, maintenance and failure diagnosis |
Make the comparison assessable
- Define the workflow, unacceptable errors and conditions for declining an answer. Score factual correctness, support from evidence and format compliance separately.
- Separate training, development and final evaluation data. Check that users, source-document families and near-duplicate cases do not cross those boundaries.
- Record model versions, prompts, training settings, retrieval corpus, generation limits and retry policy. Document any capability differences that prevent identical settings across conditions.
- Apply the same final inputs to each condition. Hide the method from reviewers and retain disagreements. Where generation varies, decide the repetition count in advance and examine variability as well as averages.
- Add source updates and access revocation as separate evaluation cases. Check paths that can return obsolete information or material the user may no longer access.
- Compare training, indexing, inference, latency and re-evaluation effort under the same assumed usage volume.
Trade-offs
Retrieval creates an information-update path while adding ingestion and search-quality operations. Fine-tuning can adjust response behavior while increasing training-data management and regression evaluation. Combining them still requires records that distinguish retrieval failures from generation failures.
Limitations
Results on one dataset do not establish a universally better method. Consider contamination, update frequency, attribution requirements and authorization. Access decisions need enforcement outside the model’s answer.
Sources
- OpenAI Model optimization
- OpenAI Retrieval guide
- What to Measure Beyond the Answer in RAG Evaluation
Related case context
These cases provide attributed design context. They do not establish that the proposed experiments or configurations were delivered in those engagements.