Y
Hacker News
new
|
ask
|
show
|
jobs
by
nh23423fefe
22 days ago
It's not in the weights. Sounds to me like jspace is the "positive cone" over relevant (large norm) j-lenses, and j-lenses are gradients wrt tokens on the residual stream when you average over some training data.