Hacker News new | ask | show | jobs
by tetris11 28 days ago
I like it! I'm just missing from the corpus for some reason...

Quick glance: TF-IDF, cosine-similarity, the only thing missing is a nice UMAP :-)

1 comments

Thanks for the UMAP suggestion. Will add.

Most of the authors are actually missing. Full processing would have yielded multi-trillion row dataset. I didn't rally have that kind of compute with me.

I have even tried running the cross-join on BigQuery... after one hour, only about 3% was done.. so, had to cancel it.

Update: Added UMAP... this is so freaking cool.. thanks for suggesting.

It's on the insights page: https://hn-buddies.stupidlabs.lol/insights under "Author Map" heading.

Super cool! I notice the projection captures the seperation better than the clustering does though. How did you cluster?