Hacker News new | ask | show | jobs
Griffin: Mixing Gated Linear Recurrences with Local Attention for Efficient LMs (arxiv.org)
6 points by vagabund 843 days ago
1 comments

"Hawk-3B exceeds the reported performance of Mamba-3B (Gu and Dao, 2023) on downstream tasks, despite being trained on half as many tokens. Griffin-7B and Griffin-14B match the performance of Llama-2 (Touvron et al., 2023) despite being trained on roughly 7 times fewer tokens."