Back to News
Advertisement
Advertisement

⚡ Community Insights

Discussion Sentiment

83% Positive

Analyzed from 376 words in the discussion.

Trending Topics

#video#model#world#robot#more#less#representation#understanding#train#disentangled

Discussion (14 Comments)Read Original on HackerNews

vessenesabout 1 hour ago
Really interesting. Upshot: a well trained multimodal video generation model has a world representation model trained inside it. They’ve done some work lifting this world model out and deploying it to robots, where it seems to work well.

On the one hand, this isn’t a new idea, and the quality video models certainly have understanding of materials, light, the world (at least in an Occam’s razor sense of understanding). I’m not aware of a video lab that’s turned itself into a robot lab yet, though; perhaps this would be a first, or a new sort of obvious-in-retrospect business path: train video model, sell video generation, scale, use scale to train robot things: profit.

I found their hands very interesting - looks like a bunch of stuff hidden in gloves - Xiami’s Robotics-1 foundation model just released videos of users training on some pretty standard looking grippers; to the point that there are demo videos of people putting on gripper type gloves to make video to train that model.

The BFL model looks like it doesn’t need that at all. Given the difficulty of the hardware side, I’ll be curious to see what they do with this.

GiffertonThe3rdabout 1 hour ago
The video at around 3.30min, where the robot arm took 3 attempts to reseat the window trim, was quite unnerving - I have not seen such resolving before. Is it new or am I way out of the loop?
dinfinity3 minutes ago
You are indeed out of the loop. Google did something arguably more impressive more than a year ago, with a VLA based bot replacing a tensioned timing belt: https://www.youtube.com/watch?v=2AAFiuEP7iE
raabout 1 hour ago
Yes that was impressive.
flufluflufluffyabout 1 hour ago
Ok the phrasing here, it’s, it’s just -

> However, compared to more specialized approaches for representation learning they produce less disentangled representations, which puts a ceiling on their usefulness for tasks that require world understanding.

Only an LLM would use a less disentangled representation of the concept “more entangled” when trying to explain to people in the real world why less disentangled representations are not as useful for modeling the real world.

pennomi14 minutes ago
Bad news, humans write silly things all the time. If anything, LLMs are less likely to make awkward phrasings than people, because they aren’t found in the training data very often.
doubleorseven7 minutes ago
this is the future. but also the present.

if you know some one out of the software development/engineering worlds, please forward this to them

rlupiabout 3 hours ago
It's nice to see partnerships between European startups.
DarkNova6about 3 hours ago
Wasn’t Flux purchased by Meta?
xnxabout 3 hours ago
takdabout 1 hour ago
Amazing work.

u.i. Zugló robot mikor?

johnbarronabout 2 hours ago
If Amazon does not buy them, the CEO should be fired...
kensai27 minutes ago
Are you American? We would prefer it to stay a European company. I personally hope Mistral buys them or somehow they unite to make a stand.
windexh8er39 minutes ago
I think you meant to type "promoted".