Back to News
Advertisement
Advertisement

⚡ Community Insights

Discussion Sentiment

59% Positive

Analyzed from 2360 words in the discussion.

Trending Topics

#don#copyright#should#model#more#companies#theft#copyrighted#fair#models

Discussion (73 Comments)Read Original on HackerNews

captainbland•about 4 hours ago
I think it would be totally incoherent to say that copyrighted information is essentially fair game to include in your model but the outputs of your model are privileged against being included in other models.

I'm not sure "competition" is necessarily the right word but I don't think that model creators and their political backers have a leg to stand on when complaining about distillation.

DonsDiscountGas•about 2 hours ago
One is allegedly a copyright violation (fair use IMHO but still a grey area legally AIUI), one is a very explicitly a violation of the terms of service you agree to when signing up. And yes I do think that if somebody puts their website behind some barrier where a user has to actively affirm they won't download the contents and use it to train an LLM that should be protected by the same law for the same reason.
hashmap•about 1 hour ago
Are you arguing that training on copyrighted material is fine but distilling against tos is somehow bad? Lmao what a take. Nah, if copyright cant even be enforced then theres no way on earth violating a tos should ever have repercussions other than they ban you. stop carrying water theyre not gonna pay you
sterlind•about 1 hour ago
many, many websites have terms of service that forbid scraping. how many of those ended up in the training set, I wonder?
FireBeyond•about 2 hours ago
So all those authors needed to do was add a TOS page in front of the contents of their books!
mindslight•14 minutes ago
It's so nice of you to think of the poor unvisited and ignored Terms of Use rather than merely clicking the nag box to proceed like everyone else. Maybe with some more positive affection the Terms might become less depressed, abstain from binge eating, and stop putting on pounds of new paragraphs every year.
zugi•about 3 hours ago
Exactly. AI companies distilled the whole internet into AI weights, now other AI companies are distilling their work to generate new AI weights. Nothing wrong with that, and the world gets open weights AI models.
Alex3917•about 3 hours ago
> I think it would be totally incoherent to say that copyrighted information is essentially fair game to include in your model but the outputs of your model are privileged against being included in other models.

IANAL, but going from copyrighted texts to an AI model likely constitutes a new original creation, not a slavish reproduction. Whereas going from one AI model to another is more likely to be considered a slavish reproduction reproduction, although arguably AI models are not copyrightable to the extent that the weights are objective facts, like entries in a phone book.

zugi•about 3 hours ago
Downloading and redistributing companies' exact AI models would violate copyright. Querying multiple AI models to form your own model with a unique set of weights does not.
captainbland•about 3 hours ago
Realistically the companies will organise their data through a network which will involve distributing and storing the copyrighted material without the consent of the copyright holder. This is in and of itself generally ruled to be a breach of copyright irrespective of whether it is fed into a model or not.
thih9•about 3 hours ago
And this stance isn’t new - AI output is already considered public domain.
SR2Z•about 3 hours ago
This is not true and there have not been any court cases which decided it.

It's entirely possible for AI output to be copyrighted if it meets the requirements; prompting an AI takes some skill.

behringer•about 3 hours ago
As of right now it's true and court decisions on other intelligences creating output is considered public domain.

https://en.wikipedia.org/wiki/Monkey_selfie_copyright_disput...

Acquiring and setting up a camera for the right shot takes some skill. That does not matter.

ncr100•about 3 hours ago
Edit: I realize I can't justify this and it's an irresponsible comment. Please ignore it.

Original:

Smells like a CEO justifying theft. It is reasonable to believe Wong thinks differently about law breaking, given this article.

Something something legal liabilities something something?

captainbland•about 3 hours ago
Which one? Seems like there's a lot of it going about.
ncr100•about 3 hours ago
I'm not sure I can defend my comment. So I'm feeling like taking it back.

Regarding CEOs - it looks pretty common yes.

I think rule breaking is a quality I've seen in entrepreneurs, thinking differently to use that euphemism. Challenging the dominant paradigm, to be tongue-in-cheek.

I imagine it's a quality engendered by business leadership schools or maybe Market economics? To Find a need and fill it, but taken to the extreme. Can be figuring out where the illegal line is and walking it.

Next, Commenting on the actual article again ..

China, the Chinese government has a known history of stealing technology from other nations. (Aside: I assume all governments do this, State-Sponsored espionage/industrial or whatever.) So that is actual theft. China is also supporting industries that it considers strategically important to success. I am ignorant at the difference between, or limitation of reach of a Chinese government supported influence of theft activities, versus a business that happens to be in China and exhibits theft behavior All on its own.

The area that I'm puzzling about is noticing how this business leader is in a position, through extraordinary concentration of wealth and power, to influence the efficacy of state-sponsored industrial espionage by helping to propose state-level policies by the United States and internationally. I don't know if the word oligarchy is the right word, but, it's like a corporation can author international financial and political policy. Pretty wild considering Nvidia was just like a $80 stock back in in the late '90s.

iloveoof•about 2 hours ago
I agree that AI companies shouldn’t be allowed to bar distillation under grounds of fairness and societal benefit.

But I don’t see how copyright has anything to do with anything. Copyright is a legal concept, not an ethical concept. Training an LLM on copyrighted material is legal. Ethically, whether a work that the LLM trained on is copyrighted or not has no relevance

dabinat•about 3 hours ago
And building a tool that can reverse-engineer other people’s products but going out of your way to prevent yours being reverse-engineered.
CodingJeebus•about 3 hours ago
The accusation of theft by anyone is really only as strong as their financial capability to enforce it. The leg that the model companies stand on is their bank account, which is increasingly becoming the only force that moves the needle in this economic and political environment.
general_reveal•about 3 hours ago
If your recommendation algorithm always recommends B when it sees A, and I derive this information from exfiltrating your masked data , is it not theft simply because I didn’t unmask the true identity of A and B?

Stole all other important information embedded in the relationship (the essence of the relationship itself).

This is not seen as theft to only two or three types of people:

1) Technically ignorant

2) Or Technically ignorant and morally bankrupt

3) Truly evil, a combination of technically capable, and morally bankrupt.

(An “Abomination” is also possible, unaware and unable to perceive morality , not just mere ignorance. For example, it is tremendously generous to call an Abomination ignorant or amoral, that’s a compliment to such a thing. Would Jensen Huang know about the words coming out of his mouth, is the true question really.)

bennettpompi1•about 4 hours ago
I think that Jensen's position as the guy selling the proverbial shovels incentivizes him to take a lot of irresponsible positions (namely around safety) but imo he's fundamentally correct here.
verdverm•about 3 hours ago
the worst thing I heard him say is that he doesn't think kids should learn their multiplication tables anymore on the Ezra podcast

while I generally agree on his open weight stances, I lost all respect in that moment, everyone of these people are so out of touch

innocent_name•about 3 hours ago
>he doesn't think kids should learn their multiplication tables

He said „The majority of the kids shouldn't learn”, not everybody. I listened to that pod and was under the impression that he wants to be portrayed as a "grounded" person and i didn't find his takes to be out of touch.

ks2048•about 3 hours ago
He comes off bad in this interview. Out of touch is correct - "I had to pump gas once a few years ago ..." and panicked because he didn't know his address or zip code.
dabinat•about 3 hours ago
Personally I think multiplication tables aren’t that useful because they’re just rote memorization. I think it’s better to teach kids to memorize a few such as x * 2, x * 5, x * 10 and teach them how to extrapolate from there. Extrapolation is a more useful skill than memorization.
nxc18•about 3 hours ago
I don’t think the rote memorization step per se is important, but the knowledge is. You can pick it up just from doing a lot of math work with those basics. Same thing in chemistry - you will learn most of the periodic table automatically just from doing so many lookups.

But IME, not having the knowledge is a major handicap. I did eventually learn essentially the full multiplication table with speedy recall and it _is_ useful in day-to-day life. I think its basic numeracy on a par with literacy if you ever have to do basics like compare prices for consumer goods, sanity check claimed facts/statistics, etc. I don’t think those skills go away when you have AI anymore than they do when you can “trust” Amazon to do the price math for you.

verdverm•about 3 hours ago
The tables have normally only gone to {12}x{12} and most young children have no issue. Why would we cut back on this?

The memorization means that you can extrapolate the more complex concepts. If you don't have the very basics, what are you even going to extrapolate on?

bobajeff•about 3 hours ago
As sometime who's memorized but also forgotten their timetables while still managing to go through Algebra I. I don't understand the importance for memorizing times tables (or any other tables for that matter).
abdullahkhalids•about 3 hours ago
Most people on a budget (i.e. the vast majority of people) are constantly using exact or approximate multiplication while shopping. And they are rarely whipping out their calculator for this purpose.
verdverm•about 3 hours ago
There are basic skills that everyone should learn, being able to do simple math in your head is a building block for more complex reasoning.

I once had a roommate who I encouraged to divide our grocery bill by 2 in his head for splitting the bill, by the end of the year he was quick and accurate.

soulofmischief•about 3 hours ago
Quick math. Basic financial literacy. Times tables are for arithmetic, the working man's math. Algebra is an entirely different math which solves different problems and deals more in the abstract.

You can definitely get past algebra without knowing your times tables, but algebra isn't going to help you in the checkout line to make sure you received the correct amount of change and bought the correct number of goods.

stevenwoo•about 3 hours ago
AFAIK he was outright lying when he said he forgot his zip code. It hasn’t been possible to buy gas in Bay Area for decades without inputting your zip code for credit card purchases. That smelled of some bad speechwriting.
rafram•about 3 hours ago
You think he drives himself and pumps his own gas? With like $200 billion in Nvidia shares to his name?
Eliah_Lakhin•about 3 hours ago
Now when the big LLM companies create their products by accumulating a notable portion of copyrighted creative works (including computer programs) from Internet, it does not count as copyright infringement or competition (or "theft"). It is considered as "fair use". So, why training LLMs on other LLMs is not fair use too?

> If you don’t like that, if you don’t like people to use your products, all you [have to do is] know your customers, and disable the service

Considering the current situation, I understand Mr. Huang point, but that's not how the copyright framework assumed to be working from the beginning. It should protect both small actors (authors) and the big companies from unrestricted use of creative works. Now this mechanism seems to be practically dysfunctional.

And a big portion of this lies on shoulders of proponents of permissive OSS, who defend an idea of (almost) unrestricted use of their source code texts for many years, and long before mass LLM scrapping became a thing.

chrsw•about 3 hours ago
Would a 100% compatible and open source CUDA stack be “competition” too?
innocent_name•about 3 hours ago
Yes, see ROCm.
HarHarVeryFunny•about 3 hours ago
I wonder how the "Chinese AI is all just distillation" folk are going to cope now that the Chinese are all aboard the post-train via agents in custom training environments wagon?

Here's Xiaomi's discussion of this, plus their open-sourcing of 7000+ RL training environments.

https://www.alphaxiv.org/abs/2609.mimo-scaling-reinforcement...

https://mimo.mi.com/docs/en-US/news/latest/v2-6

Who needs a few of someone else's "vacation postcards" of their post-training experience when your agents can go on vacation themselves!

utopiah•about 3 hours ago
Guy who sells shovels says there is gold everywhere, go figure.
RunSet•about 3 hours ago
He is selling shovels and piling manure so that he always comes out on top.
_imnothere•about 3 hours ago
It's mostly stolen goods anyway, who are they to complain about "distillation"?
prodigycorp•about 4 hours ago
Can someone answer this: if anthropic cares so much, why don't they meganerf their thinking summaries (edit: in Claude Code) the way OpenAI does? OpenAI's thinking traces so much that the Codex app doesnt even show them.
ch_sm•about 2 hours ago
In Claude Code (at least inside VSCode) thinking traces are hidden by default. You can reenable them via JSON. They’re not the real thinking traces though, they are something like thinking summaries.
aqfamnzc•about 4 hours ago
Haven't they already? (at least in the regular chat UI)
prodigycorp•about 4 hours ago
It's not there in the chatui but claude code shows more reasoning than they need to.

im not suggesting that they should, but it's so much more than the garbage both antigravity and codex show to users. seems like a thing a company paranoid about distillation would reign in.

arcanemachiner•about 3 hours ago
Pretty sure those are just summarized traces, which obscure the valuable reasoning traces used by the actual model.
prodigycorp•about 3 hours ago
yeah i know they are summarized/condensed, but they are still an order of magnitude more useful that what they competitors present to users.
Catloafdev•about 4 hours ago
Thinking traces aren't in the Claude app (at least, not visibly). I assumed it's only available (if at all) via API.
prodigycorp•about 3 hours ago
they dont show thoughts in chat, but they do in code mode.
Catloafdev•about 3 hours ago
Which version are you using? Using code mode in the desktop Claude app, I do not see thinking traces at all, and I don't see any way to toggle it.
rsx88•about 3 hours ago
Given that AI training is pure theft, I don't think this guy has the morals to talk about that.
Papazsazsa•about 4 hours ago
Competition? It's just fair play.
next_xibalba•about 4 hours ago
'Fair play' is a dimension of competition. They aren't orthogonal attributes.

You can compete fairly or unfairly.

In any event, distillation seems to be in a grey enough area that reasonable people can disagree (IMO). It strikes me as more akin to theft than to fair competition, but it doesn't seem like it should be treated as criminal. Rather, I think the onus should be on the labs to defend themselves against it.

All that said, Jensen's take is clearly quite biased.

edgyquant•about 2 hours ago
Hasn’t it already been ruled by a court that LLM outputs are not subject to copyright? It opens a can of worms to say that LLM companies can dictate what their outputs are used for.
next_xibalba•about 1 hour ago
I think copyright is mostly not relevant. Contracts operate mostly independent of copyright law (in the U.S.). If OpenAI puts in their license something to the effect of "you may not train your LLMs on these outputs" and/or "your access is limited in these ways", but you violate the license, they can and should block your access and sue you. These are things that can be monitored and enforced under existing law.

And, in fact, they already do this:

- OpenAI [1]: "[you may not] Use Output to develop models that compete with OpenAI."

- Google [2]: "You may not use the Services to develop machine learning models or related technology."

- Anthropic [3]: "[You may not use our services] to develop any products or services that compete with our Services, including to develop or train any artificial intelligence or machine learning algorithms or models or resell the Services"

[1] https://openai.com/policies/row-terms-of-use/

[2] https://policies.google.com/terms/generative-ai/archive/2023...

[3] https://www.anthropic.com/legal/consumer-terms

drannex•about 1 hour ago
Using a competitors hammer to make your own hammer is a very valid approach. That's just "free-market" capitalism for you.
Advertisement
simonw•about 3 hours ago
Jensen Huang sells GPUs to the highest bidder.
exabrial•about 2 hours ago
As it should be.
motoboi•about 3 hours ago
Jensen is borderline comical on his instances of high GPU usage. OpenClaw? "let's do it, it's amazing!", selling GPUs to china "it's of uttermost importance, USA can't afford not selling GPUs to china because other will do", chinese firms burning GPU cycles like hell on american datacenters from openai and anthropic? "this is just competition, let's do this, this is amazing".
feverzsj•about 3 hours ago
Of course. China buys shit tons of his cards for distillation.
nicce•about 3 hours ago
Expect that the best chips are still banned.
techpression•about 3 hours ago
On paper, but the world isn’t constrained by such trivial matters when enough power is behind you. Although, smuggling only goes so far, obviously, it’s not like they’re not available.

https://www.reuters.com/world/china/chinas-deepseek-trained-...

verdverm•about 4 hours ago
Sandy Monroe would likely agree, he tears down cars and sells what he learns, all the big auto companies buy info about competitor cars

Why should Ai be different?

Fortunately open weights are likely to dominate and it becomes a permissionless ecosystem

allears•about 4 hours ago
The remarks of a CEO are self-serving by definition. Otherwise he would be vulnerable to shareholder lawsuits.
Uehreka•about 3 hours ago
No. If a CEO makes self-serving comments that are in their interests, but not the company’s interests, that could be fiduciary misconduct. With that being said, CEOs have a fair amount of latitude in what they can say, since it’s difficult to tell what the “optimal” thing to say at any point would be (see: Tim Cook telling an activist investor to “get out of the stock” if he didn’t like Apple’s environmental initiatives).

The idea that CEOs’ hands are tied, and that they are kept on some sort of short legal leash, is largely (though not entirely) a fiction.

aabhay•about 4 hours ago
Don’t know why this is down voted. Jensen Huang has literally made this comment on stage when asked critical questions.

His company has been public longer than you’ve been solvent.

parineum•about 3 hours ago
Because it's a nothing statement. Everything everyone says can be categorized that way.

The shareholder part is also not correct in they way it's being implied.