|
|
|
|
|
by Nevermark
5 days ago
|
|
That isn't correct. Obviously, you can handicap any model with an indiscriminate dataset. Are you training on slop? Why? More data is not automatically better data. Just three (of many) advantages of quality data selection: (1) It gives a clearer signal. (2) It, ironically, reduces the complexity of what needs to be learned, since slop adds its own complexity. (3) It reduces dataset size which means more learning per watt (and per just about any other cost). The advantages compound. Humans are no different. Children surprise with their ability to absorb sophisticated relationships and skills, while people who had bad examples struggle to recover. 99% of human effective intelligence is higher quality cultural knowledge learned as children. None of us had to spend centuries deciding whether zero, negative numbers, the square root of -1 are numbers. Slop + quality is not quality. |
|