Hacker News new | ask | show | jobs
by rgbrgb 13 days ago
looks cool. what is latency like? I haven't used qmd before and it looks like it runs 3 local models.
1 comments

as always it depends on your hardware, the tiny models for embedding / reranking typically have low latency

qmd is focussed on local to the point of designing around single machine setups and this creates a gap where one runs agents+qmd on their laptop and LLMs on their ai box