Back to News
Advertisement
Advertisement

⚡ Community Insights

Discussion Sentiment

100% Positive

Analyzed from 84 words in the discussion.

Trending Topics

#training#text#model#vision#multimodal#opus#lot#honestly#surprised#better

Discussion (4 Comments)Read Original on HackerNews

syntaxing•about 4 hours ago
I’m honestly surprised this is better benchmark wise than the text only model. I figured the addition of vision would take away from some of the text capabilities.
mcbuilder•about 4 hours ago
One of the early results from multimodal training is that it kinda works like cross training. Training vision helps with text tasks and visa versa.
Llamamoe•about 1 hour ago
I believe that multimodal training increases the robustness of latent representations regardless of which modality is being processed.
nateb2022•about 1 hour ago
Tangentially this makes me wonder how large Opus really is. Perhaps Opus is a lot smaller than most of the 1T+ assumptions, just a lot more post-training/ finetuning on a 300-400B sized MoE model.