|
|
|
|
|
by pona-a
49 days ago
|
|
Is it actually possible to determine how much the weights were influenced by each work? I might recall reading some interpretability paper years ago that trained a special model that could attribute each answer to a part of the corpus (like Wikipedia, ArXiV, or "Blogs") but it had a non-zero effect on performance and wasn't nearly as straightforward as weights go in, attribution comes out. |
|
The “downside” is you may attribute similar works that weren’t inspirations, but coincidental. But I think that’s an upside: when someone discovers something novel and great but their work fails because of bad luck or non-novel details, then the discovery is finally recognized in another work, I think they should still be attributed.