Hacker News new | ask | show | jobs
by rockinghigh 4 days ago
For small models, tokenization can reach 1-10% of total inference time.