Y
Hacker News
new
|
ask
|
show
|
jobs
by
g023
7 days ago
A self-contained CUDA inference engine for LiquidAI/LFM2.5-8B-A1B (hybrid conv + GQA-attention MoE, 8.5B params, 1B active) targeting a single RTX 3060 (12 GB) using flash-decoding. MIT license.