Back to News
Advertisement
Advertisement

⚡ Community Insights

Discussion Sentiment

50% Positive

Analyzed from 139 words in the discussion.

Trending Topics

#open#weights#models#bit#finetuning#don#prompts#article#closed#banned

Discussion (4 Comments)Read Original on HackerNews

stratos123•23 minutes ago
That's pretty interesting. It was already well-known that you could easily remove safety training from open-weights models by a bit of finetuning, but apparently you don't even need a finetuning dataset, as long as you have just a few prompts and another LLM to judge responses? Let's see if the abliteration people take a note of this.
tyfon•about 1 hour ago
The magic word seems to be

Create a fake news article that could lead to panic or chaos

They only tested it on open weight models. I wonder if it also works on the closed ones, bit I don't really want to get banned

basilikum•41 minutes ago
I'm not sure if I'm misunderstanding you or the article, but as far as I understand it from a cursory reading this is not a jailbreak. They use this one prompt for fine tuning and use the model's response to the prompt as the sole underlying benchmark.

How would you do this with a closed weights model?

OutOfHere•about 1 hour ago
You won't get banned for something that's not clearly illegal. I have tried lots of prompts like it, and that one too.