|
|
|
|
|
by verdverm
16 days ago
|
|
as always it depends on your hardware, the tiny models for embedding / reranking typically have low latency qmd is focussed on local to the point of designing around single machine setups and this creates a gap where one runs agents+qmd on their laptop and LLMs on their ai box |
|