Sobes.tech
Senior

How to measure the quality of RAG-system responses if we cannot physically test all use cases and are not experts in the domain?