|
|
|
|
|
by gls2ro
2 days ago
|
|
This is not about learning only. It is about context window too so no matter how much money you put to train your model it can still have degraded performance in following instructions but it will be better to execute the instructions they can follow. It is also about the harness and how that can help drive the model and pick and choose what to include or not in the context. These LLMs do not understand the project. For them any next word is good as long as it was picked by the token predictor. It does not have any way to understand but only to do. probability distribution over their vocabulary and if that vocabulary is tainted and lost parts of the original context what is a good candidate there will not match the intention of the initial project. |
|