|
|
|
|
|
by limecherrysoda
3 days ago
|
|
Gemini uses MoE and context caching, which is a similar approach. You are not really accessing the biggest frontier model every time, and you're not really doing an end-to-end LLM request on each prompt. I would go so far to say frontier models have peaked and improvements from here come from clever (or very elaborate) harnessing. "LLLMHs" - Large Large Language Model Harnessing ! |
|