|
|
|
|
|
by probe
13 days ago
|
|
Do you think 3 is better than 1 & 2 as context gets larger? I think for smaller data sets its mostly fine no. It's an interesting bet by TM. End state does seem some form of continual learning (model weights update like dreaming) |
|