Hacker News new | ask | show | jobs
by ksd482 11 days ago
Grok! LOL!

Seriously, what's going on there ? Why is it so different from others? Is it just behind technologically/training wise or it's using something fundamentally different?

5 comments

Grok 4.5 is...something else.

It performs much better than composer2.5 (while being as fast). It's not Opus, but I think it's not that far off. Definitely better than sonnet for what I've been doing.

On the other hand, I think they probably heavily adapted the training data so that it really is extremely focused on code. I just recently ran my personal "poetry benchmark" on it (where I give it ~850 poems I've written over my life and ask it to comment the corpus as a whole), and it's whack. It tries to write in portuguese (most of the poems are portuguese) and code-switches constantly and mixes up words to the point of making what it writes almost unreadable (e.g. it writes stuff like "You can't QoS that that look for beast poems", in portuguese, all messed up). The quality of the analysis is also quite bad (I'd say it's definitely behind Sonnet).

So I really think they either threw away data that wasn't tied to coding so that they could fine-tune it to that, or somehow they've got such an unbalanced dataset that coding ends up dominating either way. To me, its disastrous performance in this drawing "competition" fits this narrative.

"You can't QoS that" sounds like the title of a nerdcore rap song.
Some of my poetry has clear IT jargon, but it's a very small portion of it (<1%). Some of my teenage poetry revolved a lot around the idea of wanting to become science, knowledge, and machine, and be rid of feeling altogether (to become an idea that has no body, or to be come the mathematical equations that define the world), with some very amateur odes written glorifying science and machine (a clear pastiche of Álvaro de Campos with a modern twist). But, again, this is not the majority of the work, far from it.

For some reason, OpenAI models, Gemini (and apparently Grok too), love to latch onto this and obsess over this idea that it's "programming poetry" or "poetry for the IT crowd". Often OpenAI and Gemini try to write the "equations of my poetry" (granted, I do write about a cyclical relationship between thinking, feeling and writing a lot, and I do have ONE poem which ends with a Q.E.D.).

I'm giving this context to say that it is very bizarre. It's as if they latch onto it and act as if it's a core or highly distinguished part of the poetry, when it really isn't. Anthropic models, on the other hand, absolutely do not do this, and have never done it.

I really don't understand why this happens. Maybe it's because it has a lot of portuguese, I don't know. And even though the "QoS" is clearly the wrong token being generated, I have had situations where gemini spoke of some phase of my poetry as the "Q&A part" (really, no joke...)

In any case, it's why it's my personal benchmark after all :D

might also be U shaped memory issue did you try changing the order of your poems to see if it focusses on different ones? maybe it just happened to have those old ones in points it memory was focused at. (ofc its not good but would be an alternative idea to it focussing on IT things/code to produce such results)
Yep, I did. It didn’t have any effect. Randomizing the poems has stopped having a significant effect in the last year or so on the analysis as a whole.

It still somewhat affects the way the AI looks at the poems, especially if it has to make lists of the “best”, but it is not a very pronounced effect nowadays. Maybe it gives some preference for earlier and latter poems but not a lot.

Back when I started doing it, these holes in the context were super obvious, just like you described. It would mixup entire periods and also hallucinate or forget about specific sections. That’s precisely why I started randomizing them as part of my experiments.

Nowadays the frontier models are much better and don’t really mix anything up.

They focused a little too much on Grok Imagine.
Did Elon tell some poor engineer to give grok a prompt injection for drawing,"make it look like one of my childhood drawings!" Just like the Tesla truck?
The razor-wire at the bottom for Starry Night was clever, and very Grok. Really shows its military spirit.

Edit: I just don't see the point of redacting the Mona Lisa

I read the balls as “houses”, though the phallic spire emerging from them dead-center is also very on-brand
I must admit, I have a natural tendency to overlook C&Bs, but solid catch -- it's there. Perhaps that explains the Mona Lisa redaction; Grok probably put more effort into that one.
Redacting? Is that not just the model using the smudge tool it was given.
Is this not someone being near infinitely obtuse?

HN, the only environment in the world where people will scientifically falsify a joke.

If I saw something I thought resembled a joke I wouldn't have made the comment.
ESL?

Allow me to explain the theory behind the humor...

Grok's art had heavy use of what strongly resembles a permanent marker with lateral-ish scribbles similar to what one might see when someone is trying to mark something as discarded. Documents, especially those procured by the US government when responding to FOIA requests, contain many 'strike-through' style black marks resembling bold permanent-marker lines, which obfuscate the words they wish the reader to remain ignorant of. In many cases, this is the entire document, especially where unredacted words might otherwise serve in any way to clarify the nature of the FOIA inquiry, i.e. answer the question.

Grok's art, resembled to me, and others potentially appreciating the joke, redactions more than art. In some forms of humor, the presenter of the joke feigns obtuseness, or a mildly self-deprecating, stupid perspective. But the stupidity of the perspective actually reveals a relevant concept, or idea.

Often, such jokes can be deciphered using common sense, or various razors, including both occam and hanlon, the former self-suggesting that the joker is probably not so stupid as to think Grok was really redacting anything, and the latter self-suggesting that there was probably no malice involved.

It doesn't get any more funny when you try to explain something that wasn't funny in the first place.

Hence, I can't be bothered to read any of this.

the starry night one is soo funny