Hacker News new | ask | show | jobs
by bigmadshoe 14 days ago
Good point! I thought you meant splitting them and then doing inference with some kind of learned router while keeping all the split models loaded at once. What you're suggesting is pretty sensible.