
Assistants that answer from your own documents
For an internal knowledge assistant, start with organized, current, permission-tagged documents and tested retrieval—not a model comparison. Model choice still matters, but no vendor-neutral evidence establishes that it matters less than document organization in every case. The practical priority is to make the knowledge base retrievable and trustworthy, then test models against the questions your staff actually ask.
How does an assistant answer from your documents?
A document assistant usually does not retrain a language model on your files. It uses a process commonly called retrieval-augmented generation, or RAG: find passages related to the question, add those passages to the model’s input, and ask the model to compose an answer from that material. This allows it to use private or frequently revised information without retraining the underlying model. That basic sequence is described by Microsoft, AWS, and the original RAG paper.
In practical terms, someone asks, “What is our approval process for expenses over the limit?” The system searches the indexed material, retrieves the passages it considers relevant, and supplies those passages with the question. The model writes an answer and, if the system is designed to do so, returns links or citations to the source material. When the source changes and the index is updated, later answers can reflect the revised information rather than a model’s older training data. Microsoft’s RAG guidance describes these grounded-answer and citation uses.
This distinction matters because the model does not acquire a durable, complete understanding of the company handbook. It receives selected pieces for each question. If the right passage was never ingested, was parsed incorrectly, was blocked by a filter, or did not rank highly enough to be retrieved, the model does not have that evidence. Microsoft explicitly notes that irrelevant or incomplete retrieval can still lead to incomplete or inaccurate answers.
What can it do, and what can’t it do?
A well-scoped assistant is useful for finding and explaining material that already exists. It is not a substitute for missing policy, conflicting instructions, or professional judgment. The boundary is easier to see when the intended job is written down before the build starts.
| Job | What to expect | Important condition |
|---|---|---|
| Answer factual questions from policies, manuals, FAQs, and repositories | A suitable use | The relevant passage must be present, indexed, and retrieved. Source |
| Summarize retrieved passages | A suitable use | The summary is limited by the material retrieval supplied. Source |
| Show citations | Useful for traceability | A citation does not by itself prove that every generated claim is supported by the cited passage; citation correctness has to be tested separately. |
| Reflect a revised policy | Possible after updating the source and index | There must be a working update process. Source |
| Guarantee that every answer is true | Not reliable | Grounding reduces dependence on unsupported memory but does not eliminate retrieval or generation errors. Source |
| Repair missing, outdated, or contradictory policy | Not reliable | The source owners must resolve the underlying knowledge problem. Source |
| Make compliance judgments or compare many complete records | Do not assign this to a basic assistant | These tasks generally need additional retrieval, deterministic rules, specialist review, or a purpose-built workflow. Source |
The ability to say “not found” is part of a dependable design. Retrieving more passages does not automatically make an answer safer. Research has documented a distracting effect in which related but irrelevant material steers models toward wrong answers, supporting selective retrieval rather than filling the input with every possible match. An ACL 2025 study examined hard distracting passages, while earlier peer-reviewed work also reported the effect of irrelevant context.
Why does document organization affect answer quality?
Retrieval works with what the documents make explicit. Descriptive headings, defined abbreviations, section summaries, consistent terminology, and smaller self-contained documents give the search process clearer material to match. Ambiguity, duplication, insufficient domain context, and missing metadata make that job harder. AWS lists these as common retrieval problems and recommends those structural improvements in its discussion of RAG document challenges and writing best practices.
That guidance is implementation guidance drawn from enterprise experience, not a controlled measurement proving that each formatting change adds a fixed amount of accuracy. AWS describes the recommendations as curated from experience with organizational leaders and enterprise use cases in its next steps. There is no universal best chunk size, number of retrieved passages, embedding model, or reranker in the supplied evidence. Those settings must be tested against the particular documents and questions.
Layout-heavy material needs separate attention. A table can encode meaning through rows, columns, merged cells, and position; a diagram can put the important relationship in the image rather than its surrounding text. Plain text extraction can lose those relationships. Layout-aware parsing, multimodal processing, or a written restatement may be needed, according to guidance on RAG pipelines for PDFs and document preparation.
The actual performance on a company’s scans, tables, diagrams, spreadsheets, or handwritten material is not known until representative files are examined. That changes the build order: test the awkward documents early instead of demonstrating the system only on clean policy pages.
Does organization matter more than model choice?
Not as a universal rule. In our assessment, document organization and retrieval are usually the most controllable starting points because changing the model cannot restore a passage that was omitted, misparsed, or hidden by incorrect metadata. But the model remains an independent part of answer quality, especially as inputs become longer or answers require combining evidence from several documents.
| Research result | Scope and date | What it means here |
|---|---|---|
| 2,326 test cases across 4 question-answering categories, 3 kinds of naturally occurring long text, and 11 language models | LaRA benchmark, as of February 14, 2025 | The best approach depended on model size, long-text capability, task, context, and retrieved chunks. This is benchmark composition, not an expected accuracy rate for company documents. |
| 13.9%–85% performance degradation as input length increased, despite perfect evidence retrieval | Study of 5 open- and closed-source models, as of November 2025 | Perfect retrieval did not remove model-side difficulty. The range spans mathematics, question-answering, and coding tasks and is not a forecast for an internal assistant. |
| Up to 20% performance loss as document count increased while total context length and evidence position stayed constant | Multi-document study, as of November 2025 | Most tested models declined, but Qwen2 remained consistent. That exception directly shows that model choice can matter. The custom datasets came from a multi-hop question-answering task, not a company knowledge base. |
| 3 language models in a RAG-versus-long-context comparison | Comparison published as of July 23, 2024 | Long-context processing performed better on average when sufficient resources were available, while RAG retained a comparative cost advantage. The paper predates newer models and did not establish a universal currency cost. |
The honest conclusion is conditional. Clean documents do not make all models equivalent, and a stronger model does not make weak sources dependable. A representative evaluation must test the complete chain: source quality, parsing, retrieval, permissions, context assembly, answer generation, and citations. No universal accuracy rate is available because results depend on the organization’s documents, questions, permissions, retrieval setup, and model.
When should you use RAG instead of the complete document?
RAG is a reasonable choice when the collection is large, changes often, or needs metadata and permission filters. It lets the system select a smaller set of relevant passages instead of supplying the whole repository with every question. Metadata such as dates, departments, projects, document types, and classifications can support relevance filters, freshness rules, and permission-aware retrieval, as described by AWS and Microsoft.
RAG is not automatically the best choice for a small collection or a question about one bounded document. Comparative research found that sufficiently capable long-context models can outperform conventional chunk-based RAG on some question-answering benchmarks. For those cases, supplying the complete relevant document—or routing different questions between full-context and retrieval methods—deserves a direct test. The evidence comes from a 2024 comparison and a later study of long-context and RAG approaches.
This is not a platform-selection article. The workflow-tool tradeoffs belong in “n8n vs Make vs Zapier — and why we don’t have a favourite.” Costs also depend on document volume, update frequency, query volume, model, hosting region, and security requirements; that work belongs in “What does it actually cost to automate a process?”
How should permissions and untrusted documents be handled?
Permissions must be enforced when passages are retrieved, not only when links are displayed. Otherwise, the assistant can use a restricted passage to compose an answer for someone who was not authorized to read the source. Microsoft’s retrieval hygiene guidance, its RAG overview, and the AWS security reference architecture all address access controls around retrieved context.
Documents are also inputs, not trusted instructions. A compromised or malicious file can contain text intended to manipulate the model through indirect prompt injection. Source provenance, ingestion controls, access restrictions, sanitization, logging, and adversarial testing are documented defenses in Microsoft’s RAG prompt-engineering guidance and the NIST preliminary report on adversarial machine learning.
If existing source systems cannot expose sufficiently detailed permission and change metadata, do not assume the assistant can reproduce their access rules. Confirm that before broad deployment. The wider question of service providers, workflow access, and data handling belongs in “Who can see your data inside an automated workflow.”
How do you test whether the assistant is good enough?
Build an evaluation set from representative staff questions. Include straightforward lookups, questions that require two passages, outdated terms people still use, restricted documents, contradictory sources, questions with no answer, and files with difficult layouts. Then inspect whether the system retrieved the right material, respected permissions, declined when evidence was missing, and produced claims supported by its citations.
This evaluation is not optional maintenance work created by a bad implementation. RAG itself adds a knowledge-management lifecycle: documents must be ingested, parsed, cleaned, divided into retrievable units, indexed, permissioned, updated, and evaluated. AWS describes these stages in its RAG overview and data lifecycle guidance.
Set ownership as clearly as the retrieval settings. Someone must decide which policy is current, what replaces an obsolete document, how conflicts are resolved, and when an index update is complete. If nobody owns those decisions, the assistant exposes the disorder more quickly; it does not resolve it.
When is a document assistant the wrong answer?
Do not start with an assistant when the source material is mostly absent, when contradictory policies have no owner, or when access boundaries cannot be enforced. Fix the knowledge and permission problem first. Also do not present a basic assistant as a dependable compliance judge or as a system for complex comparison across long, unstructured records; Microsoft’s implementation guidance and research on long-document reasoning support adding specialized workflows and review for those jobs.
If the process needs consequential decisions rather than document lookup, read “Where AI agents go wrong — and where to put the human.” For the broader decision not to automate, see “When automation is the wrong answer.”
What should you do next?
Choose a narrow body of documents and collect representative questions, including failures and permission edge cases. That is enough to test whether the sources, retrieval, and model can support the actual work before expanding the scope.
Book the call. Bring a sample of the documents, the questions people ask, and the access rules that must survive the build.
