We built a live demo to compare DeepSeek R1 and OpenAI’s o1 model on a simple RAG pipeline using Haystack, our open-source AI framework. Both pipelines use the same retrieval and prompts to ensure a fair comparison.
Would love to hear your thoughts -- how do open-weight models stack up against proprietary ones for RAG?