|
|
|
|
|
by xg15
2 hours ago
|
|
> I gave it the first paragraph of the Lord of the Rings, a 1M token budget (~$10) and asked for three js render of it. I think it's interesting that the "Bag's End" interpretation in the video clearly looks like the one from the movies, but generated here as a three.js 3D asset. It makes sense that the movies (or shots/frames from them) were in the training data, and I can also easily imagine an association in concept space between the textual description of Bag's End and the frames from the movie. But how on earth does the model then go on and convert the latent representation of those images into coordinates for a 3D mesh, without ever even restoring the image? In what kind of representation are the images from the movies stored that it can do that? |
|