| HN Mirror

Y	Hacker News new \| ask \| show \| jobs

by giang_at_glai 112 days ago

Author here.

This post shows “concept algebra” on language model: inject, suppress, and compose human-understandable concepts at inference time (no retraining, no prompt engineering).

There’s an interactive demo on the post.

Would love feedback on: (1) what steering tasks you’d benchmark, (2) failure cases you’d want to see, (3) whether this kind of compositional control is useful in real products.

2 comments

anon291 111 days ago

I would personally like some quantification of how good this is compared to just replacing the system prompt of an off the shelf 8B parameter language model.

The suppression bit is very powerful. I would like to see a quantification of how often a steered 'normal' language model will mention things you asked it to suppress vs how often this one does

link

giang_at_glai 111 days ago

We will share a technical write-up soon that addresses both of your questions: (1) steering vs. prompt engineering, and (2) how effectively our steering suppresses undesired generations.

If you have joined our waitlist, we will notify you as soon as it is available.

link

didgeoridoo 111 days ago

Hi! Have you published the concept dictionary yet? I’m looking into using Steerling to investigate how different moral scenarios elicit various responses in LLMs (using Haidt MFT concepts mostly), and my first few inference runs have been hamstrung by not having a canonical mapping of concepts to IDs. Thanks!

link

luulinh90s 111 days ago

Hi! Thanks for checking.

We haven’t published the concept dictionary yet.

We plan to release it in soon with other important artifacts.

link

didgeoridoo 110 days ago

On the waitlist — please announce it there!

link