Hacker News new | ask | show | jobs
by yiyingzhang 2 days ago
This approach only works for small context requests. For large context and relatively smaller output (say understanding a huge code base), the cost will mainly be on prefill, and sending the large context to multiple models will only increase the cost, possibly by some factor.