Advertisement
Advertisement
β‘ Community Insights
Discussion Sentiment
50% Positive
Analyzed from 55 words in the discussion.
Trending Topics
#models#post#training#inherently#interesting#claim#frontier#fact#saturated#gets
Discussion Sentiment
Analyzed from 55 words in the discussion.
Trending Topics
Discussion (1 Comments)Read Original on HackerNews
Gets me thinking whether this means that post-training inherently has a hard time to get rid of facts. That is, on a general corpus, can we reasonably prevent models from knowing/outputting the things they're not supposed to via post-training, or does that only add obstacles to the recall? Are all models inherently jailbreakable?