Aurornis warns removing refusals doesn't add missing knowledge
4 Sep 21 9:52 AM · 6d ago · 3 comments · 1 source · development 4 of 5
A technical comment argues that because training data is often shaped around refusals, unlocked models may lack the underlying knowledge to answer correctly and can instead produce confident hallucinations, with possible quality drops on unrelated questions.
“If the model was trained on data that gives a refusal to that topic, the real information might not be encoded in the model at all.”
AurornisHeretic project Open-source tool for removing refusal behavior from LLMsAurornis HN commenterTepix HN commentermatheusmoreira HN commenterTristanDaCunha HN commenter
The whole story articlespostscomments the bright band is this development · numbered dots are the others · click one to jump
What people said 4 voices · best of 5 · verbatim
-
Two problems with modifying models like these, which you should be aware of.First, the training sets of these models are usually shaped around the refusal, too. They might not have enough of the knowledge to answer correctly even if you stop it from going down the refusal path. If the model was trained on data that gives a refusal to that topic…
-
This is off topic.Why don’t people who release python projects ever encode the venv steps into the installer? Can’t pip just do that step for the user?
-
Does this actually modify the weights?It submits prompts that get refused, then detects and modifies the weights responsible?Like brain surgery?
-
These "safeguards" are actively contributing to computer insecurity at this point.
All 5 developments of "Heretic" tool for stripping refusals from open-weight LLMs… →
Hacker NewsMastodonNewswires