The idea
A language model knows a lot about the world and nothing about your business. RAG, short for retrieval-augmented generation, gets round that by searching your documents first and handing the best passages to the model as it writes its answer.
Because the answer is built from those passages, it can point back to them.
Why not train it on our documents?
You can fine-tune a model on your own material, but it is slower, dearer and goes out of date the moment a policy changes. With RAG, a corrected document is available as soon as it is indexed, and access rules can be applied at the moment somebody asks a question.
Where it goes wrong
Almost always in the same few places: old or contradictory documents, search that misses the right passage, and permissions nobody thought through.
A good system cites its sources, is allowed to say it cannot find something, and is tested against questions where you already know the answer. The model matters less than people expect. Most of the work is in the documents and the search.