ZH version is available. Content is displayed in original English for accuracy.
Advertisement
Advertisement
⚡ Community Insights
Discussion Sentiment
51% Positive
Analyzed from 5084 words in the discussion.
Trending Topics
#copyright#human#generated#code#content#copyrighted#don#should#case#protection

Discussion (144 Comments)Read Original on HackerNews
https://en.wikipedia.org/wiki/Monkey_selfie_copyright_disput...
Hopefully no judge uses that as precedence.
I'm actually having a bit of trouble thinking of what sufficient societal good there is/would be in granting copyright on raw dashcam or security camera or the like footage? None of those purely mechanical automated systems need a subsidy or encouragement to generate more. Certainly someone can use that sort of thing in the creation of a copyrighted work but what would be the issue with the underlying material in that case being unprotected?
Details https://ftp5.gwdg.de/pub/gnu/www/philosophy/words-to-avoid.h...
In any case, any country that decides neural network generated media isn't copyrighted then faces a major, and probably impossible, problem of trying to prove it wasn't made by a human.
For what it’s worth, dashcam footage can absolutely be copyrightable.
The ruling is more about “only humans can get copyright protection”, not so much anything about whether a button is pressed or not.
What if this applied in a photography class? The instructor owns the equipment and helped “set up” the photo. Does the instructor own the copyright?
It's fine, I guess. How does it work in cinema? A director who is the creator of the project must have to get rights from every camera/mic operator.
So you're more of trying to create a special rule where if the normal recipient of a copyright would be invalid, then it slides to the 'nearest' most appropriate individual, but that seems extremely fragile and difficult to define.
It's like things that are already in public domain. Even if you make a coloring book out of paintings in public domain, it doesn't necessarily mean others can just print your book as-is.
Have you been involved in copyright or patent litigation?
Because it's not that easy.
If you're on the defendant side of a copyright violation case, it's extremely hard to use "well the original author didn't really make it...* as a defense. (Patent cases are often defended with this argument though, as a patent grants far boarder protection than copyright and can be rejected on prior art. But still it's very different from "AI made this actually.")
If companies get slope with creative output to the point that "a few employees" can reproduce it in shadow markets, don't expect to get copyright protection without giving governments revenue and speech-control.
If that is faithful reading of the law, that makes sense. I know a number of people who use AI, but none of them (that are making anything actually useful) have the output "entirely generated" (aside from some POC tests that never see the light of day).
I have a hard time believing anything of value, anything worth copyrighting, could be entirely generated by AI.
Perhaps you're not stretching your imagination enough. What if an expert novelist used an AI like it were a fancy auto-completing dictation machine to write the next great American novel? AI may have "entirely generated" all the text, but what if they micromanaged the shit out it?
I can imagine the difference between someone who fires off a lazy 5 minute prompt, and someone who labors for months and months to get exactly the results they want.
If companies fail to protect their investments in generating IP, they will stop investing in generating it.
And unless IP generation costs (all in, including the humans telling them what to generate) fall close to zero, it will be bad for the world if companies cannot recoup investments in generating new IP.
We would expect this to hit those industries relying on IP protections the most, e.g. pharma.
Take note of the qualifier entirely. If you're working with an agent steering it to produce the results you want, it would be an entirely different story.
Which verb would describe their knowledge of this security code?
However, there's no issue with including non-copyrightable code in otherwise copyrighted projects. There's already plenty of non-copyrightable code like auto-generated boilerplate.
That is not how copyright/trademark/contract laws work, and isomorphic plagiarism is not a long-term business model. People also loved Napster at first too. Good luck =3
https://www.youtube.com/watch?v=YhgYMH6n004
So no license is enforceable with code written by AI.
The issue is most GPL license fall under contract law, and scraped code can't legally have assigned "copy" rights on an "AI" vector search compaction output.
https://www.youtube.com/watch?v=YhgYMH6n004
Indeed, these rules obviously don't apply in places like India, Russia, Iran, and China. =3
As the dark specter of Disney Mickey Mouse looms over every LLM model involved in isomorphic and character plagiarism. Yes, even motion capture is considered a performance act in the guilds, so video reskinning an unlicensed performance act people make is also a liability.
It would sure save a lot of money if you don't get caught, so people are gonna try it for sure. =3
Perhaps we'll have new iterations of FOSS licenses to adjust to legal declarations.
All models know what Disney Mickey Mouse looks like too. =3
They have a copyright on the original, and if they hire a human, the human would have a copyright on the translation (which would generally be licensed or transferred back to the author in some way).
If they use an AI for the translation, by the logic here, the translation wouldn't have its own independent copyright, but (based on other long established principles of copyright) it would still be a derived work of the original, so even if this decision holds it would not be legal to make unauthorised AI translations, pirate authorised AI translations, make further translations into other languages (or back to English), etc.
Which seems fairly reasonable! But consider:
If you start with, say, a 90,000 word novel, and ask for a translated novel, you (presumably) have sufficient rights to stop someone making unauthorised copies of the AI translated version.
If you start with a 300 word prompt, and ask for a logo, you (apparently) do not have sufficient rights to stop someone from using it without authorisation.
So some combination of the input (0.3k vs 90k) and the output (logo versus novel) crosses an inflection point between these two extremes, and I think it's interesting to wonder what the boundaries are. Like, in theory you could graph input size versus output complexity, and sketch a frontier between "the author's protected expression survives in the output" and "the author's protected expression does not survive in the output". And I don't have the slightest idea what I think a fair frontier would look like.
It has always been like that though. Copyright is just that arbitrary. You really can't tell if Android violates Oracle's copyright over Java by reading law text.
Monkeys and LLMs (so far) either don't understand or don't need such incentives, so they don't get copyright.
But it just begs the actual question of how much human contribution there needs to be:
- I wrote the prompt (not enough)
- I wrote many prompts and iteratively refined them using distinctly human skill (open question, but loosely seems still not enough, potentially in the EU but maybe in the US?)
- I made minor modifications post-generation (open question, probably enough)
- I made equal or more contribution to the final result (this better clearly have copyright protection or we are in real trouble)
Hmm I don't think that would be enough. I'd expect that would only make the modifications themselves copyrightable, but not the whole modified work including the AI parts.
Compare for example the case where the US copyright office ruled that, when assembling AI-generated images and human-written text into a comic book, only the human-made elements themselves (text, arrangement) get copyright protection, but not the images.
E.g. if you generate an AI photo and color grade it, i expect only the color grading would be protected (if that is even significant enough to be protectable), not the rest. And someone else could re-color-grade the same image without infringing your copyright.
Some examples of why I think this is really hard: Say I build a story generation system. I work hard on building an agent swarm of actors, critics, editors, researchers. I craft into the various agents concepts of story arcs, outlining techniques, character development. I build a huge well thought out process for how to agentically write an actually good story, so long as you give it a title. Heck, I even design and train my own custom LLM with original layer ideas and novel training techniques to use on this system. After all that I then take that final step and give it a title. Do I have no claim to that? I probably put more work and creativity into it than an author would have a book. What if I then gave it 500 titles? 5,000? Would my claim degrade the more titles I fed it? Is it a percentage of work question? What is the core concept here that defines 'human-centric'? What is the cut-off here?
Let's go even further. I don't prompt. I live in a world with unlimited context models. I have a conversation about the book I want it to write. During that process I reject some ideas and accept others. I didn't give it a 'system prompt' but essentially all I did was prompt it and select versions I liked. Is that not human centric? How about if I asked it for advice and it did some editing work on my story? Did that make it not human centric even though the starting text was mine? What if that starting text was 99% replaced with a version 10x as verbose. Defining based on how you interacted with the model (prompted and selected) just seems way to weak to be a clear test.
If we go back 10 years and your friend says “I have an idea for an app, here it is,” and you build it, you own the copyright because you wrote it.
You give an idea to the pile of math calculated of the stolen work of humanity, the math owns it (which it can’t, so no one owns it).
No matter how detailed of a conversation you have with a friend, I don’t think they have justification to claim copyright over code written by you.
Based on your logic that should not qualify, but it currently clearly does: https://quayola.com/selected-unfinished-sculptures/
But if the tool is created from collective human creation, the copyright should belong to all humanity, not the person who triggered the tool.
If you trained an LLM entirely on your own input, I think you should own the output, but that is not the case for any widely-used llm.
It's an antiquated mechanism and is far more abused than it is actually used at this point.
Does it enable decompilation remasters of classic games?
It feels like AI is a cleanroom laundromat
2) If one applies a copyright message to AI generated output, is that fraudulent?
The AI system must function merely as a tool or instrument (like a camera or Photoshop) guided by the human, rather than acting as the creator itself. The line may get a bit fuzzy case-by-case, but effectively the human must be the creative one, not the AI.
This is not unprecedented. Machine generated technical data, sensor outputs, automated surveilance photography, monkey selfies, purely algorithmic or generative music and such were already disqualified long before AI came along.
The set of possible and desirable non-copyrighted text/audio/visual states to render is effectively infinite. Laws don't prevent Scrabble clones; tropes are not protected.
Endless remixes of public domain content are an option as well.
If an iPhone with model weights in chip, a local Mac mini with similar but more powerful model on chip ends up capable of generating endless content copyrights won't provide a moat.
https://fmhy.net
Yes, that is vague. I think the examples were like:
If you paint a symbol and use an AI filter over that to stylize it, you own the symbol aspect of the image but not the stylized final result.
You can own a book of AI images as a curated collection. But, not the individual images.
It feels shitty that they have been trained on the life sums of all of our work and online presences with absolutely no credit given... But then again, I'm not sure I'd want to know what parts of the weights were from me and which weren't.
Say I would prompt Unix like system in few prompts or started agent chain to make it. That OS wouldn't have a protection.
He became a bit of a pop phenomenon, with stores dedicated to selling high quality prints of his works.
And one of the things, so I'm told, that could happen, was that you could buy one of his prints, but someone at the store, an artist, would dab some paint onto the print. Add some "light" to it.
Obviously this is a commercial endeavor, so I doubt the provenance of the rights holder was in doubt (such as through an employment agreement).
But it's, perhaps, an interesting case study about who owns what in a time of augmented and manipulated media.
entirely is the plank supporting this.
The creativity requirements may seem arbitrary but there’s a legal distinction between a sculpture and a standard brick.
Or more relevantly, a recipe find on recipe sites (with the author's entire backstory) vs a sequence of instructions. The latter is not copyrightable, even if there was some creativity that went into it (eg. word choice).
In fact, the defense of "I wrote the prompts that led to the code that the LLM wrote" would be a much better defense.
However! Since AI work is noncopyrightable, the AI’s effort in translation is simply ignored. Claude’s The Odyssey would have zero rightsholders, and remain public domain. ChatGPT’s translation of The Three Body Problem would still be under Liu Cixin’s copyright.
AI only = no copyright.
“A Single Piece of American Cheese” got a copyright because it had human involvement in compositing.
Theatre D’Opera did not because it was primarily prompt driven.
Thaler didn’t because he said it was machine derived.
Humans must be involved for a copyright.
If the AI is treated as an agent during inference, distinct from its user and not merely as a tool, then this should also apply during training too.
Based on this, it seems that AI agents are consuming people's code without permission. The MIT license only gives rights to "any person obtaining a copy of this software".
So the rights are given to a 'person', and the rights pertain specifically to a person who performed the act of 'obtaining a copy of this software'.
MIT license says 'obtaining a copy' and uses the word 'software', not 'code'. 'Software' is to 'code' what 'shop' is to 'building'; if you bought the shop, it doesn't necessarily mean you own the building. These are two different things and require different clauses. MIT explicitly separates the two and emphasizes that the author of the software retains copyrights (presumably over the code as this is the only thing over which they could claim copyright).
The code is different from the software; you can write the exact same software which behaves in the exact same way using completely different code; can be poorly written or well written. The difference is extremely meaningful to the person who invested effort to write the code in a clean way.
If we say that the agent is a separate entity from the person who ran it during inference, surely the same distinction can be made concerning the person who ran the agent during training. So the term 'person' from the MIT clause doesn't seem to apply here since the person running the agent has been factored out (just as they were during inference). Also, the agent is not obtaining a copy of the software; it's obtaining copies of the code which is copyright and independent of the software (as the MIT license clearly asserts).
Copyright didn't always exist, nor should it continue to. Hell; it must not.
I think the words (read: hilarious 1.25pp pamphlet) of Aaron Swartz on the topic are just too poignant to ignore, given the paths of Reddit (corrupted yet democratic), IP law (malignant yet showing cracks), and government survellience have taken in the Trump era. Despite the dated context... he really says it best:
https://ia800101.us.archive.org/1/items/GuerillaOpenAccessMa...
I mean the copyright has to belong to somebody right?
Why would it? It’s generated by blending together ~ every bit of content on the internet and in books that they could steal. Why would the operator of the blending machine suddenly get copyright?
Does a gambler own the copyright on the symbols generated by a slot machine?
If you record yourself reading a book, you own the audio recording copyright but it would be a copyright violation to reproduce that copy without a license for the underlying rights.
In this situation:
If the AI generates the audio recording of a book, no one owns the copyright of the audio recording but it would still be a copyright violation to reproduce that copy without the underlying rights.
What in the world makes you think that?
What is this implication based upon? Where does it say in copyright law that if no copyright arises for your work then it does not infringe the copyright holders' rights? This is completely devoid of logic...