Hacker News new | ask | show | jobs
by imperio59 10 days ago
Pre-training data is pre-tokenized ahead of time before being used to not waste any GPU compute.

A massive speedup like this is a nice efficiency savings on some of these data pipelines for sure.