Hacker News new | ask | show | jobs
Basaltlabs Monolith-1.0 – #1 on Last Exam, AIME, GPQA Diamond, MMLU-Pro (basaltlabs.org)
1 points by Topfi 10 days ago
1 comments

Generally, I am very sceptical when it comes to public eval results, as I have in the past seen labs verifiably train for specific benchmarks [0].

Having read the technical report [1] however, I am absolutely certain that this is not the case here as they took such great lengths to prevent any training data contamination and due to this I bestow upon Monolith-1.0 the honour of being the first model I have such faith in that I do not see a reason to run my own benchmark suite. It is clearly Super AGI.

Here is a Pelican on a Bicycle to assess that aspect of the model, seems slightly ahead of Geminis output to me: https://imgur.com/a/9brUIow

/s for clarity.

Here the video [2] that covers how they did this along with the insanity that is the uncritical social media reporting and ease of creating hype via benchmark numbers, something repeatedly and provably used to great effect by some actual labs. Labs have done this and they will continue doing this as long as the incentives outweigh those trying to hold them to account.

[0] https://news.ycombinator.com/item?id=48951229

[1] https://basaltlabs.org/Monolith-1.0-2606.pdf

[2] https://www.youtube.com/watch?v=enk4w5mRjQY

Because the previous comment doesn't quite spell it out: this was an attention seeking hoax.

"I Faked the World's Best AI Model"

https://huggingface.co/basaltlabsai/monolith-1.0 has been pulled down with the message:

"The experiment for this has concluded - you can watch the YouTube video for an explanation: https://youtu.be/enk4w5mRjQY

Given that the model intended for the public is an inflated version of the original Qwen 2.5 7B Instruct model (https://huggingface.co/feisah65/5-v6-exp), and given that this repo exceeds Hugging Face's free tier limits and provides no value, the weights of the ~3TB model have been removed."