RAG systems often look straightforward in a prototype, but production quality depends on much more than model selection.
In our experience, the main variables are document preparation, retrieval quality, metadata filtering, access control, citation support and a repeatable evaluation set. Monitoring failed retrievals is also more useful than tracking only generic model metrics.
We build these systems at [url=https://ai-development-services.com/]AI Development Services[/url]. Which retrieval or evaluation metric has been most useful in your environment?