Y
Hacker News
new
|
ask
|
show
|
jobs
by
gpjt
195 days ago
OP here: one thing that surprised me in this experiment was that the model trained on the
more
curated FineWeb-Edu dataset was worse than the one trained on FineWeb. That is very counterintuitive to me.