> there is no value in being a Public-Domain-Dumping-Ground.
First of all, isn't that the entire point of the Internet Archive? Sounds pretty valuable to me.
Second, doesn't the same issue effectively apply to everything written with the BSD or MIT license? I'm not sure I understand the point though, so I'm probably missing something.
Did you believe that the Internet Archive hosts nothing but PD content?
The Internet Archive has lots of PD books and other works that have been scanned in. But the Internet Archive's web pages are all copyrighted. Only a small percentage of anything on the Internet is PD. All those news articles that HN bypasses paywall? Piracy of copyright.
How do you think the IA got in legal trouble during the pandemic lockdowns?
There are instances of purely PD archives: librivox.org and gutenberg.org. Their mandate means everything is PD, whether it was PD to begin, or whether it is novel work by modern creators. The archives are sustained by their donors and volunteers, who are scanning these works, fixing up the OCR, recording audiobooks, and otherwise donating their time, talent, and treasure to building the archives. Perhaps both are obsolete by now.
Everything written in BSD and MIT license is copyrighted. Copyright may be assigned, dual-license status may exist, non-exclusive rights may be assigned. None of this applies with purely PD works.
I stated no mere opinions, and I made no novel predictions, but I gave observations of current practices. The only prediction is that these two practices expand and escalate for as long as LLMs' legal status remains this way.
A BSD or MIT (or any other) license doesn’t remove copyright. For example, copyright reserves the author the right to also release the work under a different license in the future. AI output doesn’t have such associated rights.
1. The copyright claim is valid in places where LLM output is always covered by copyright.
2. It may be valid in places where sufficient human contribution makes AI code covered by copyright.
Even if it is not covered by copyright everywhere, it has a significant effect. If something is covered by copyright in some countries but not others its cannot be globally distributed without a license.
> 1. The copyright claim is valid in places where LLM output is always covered by copyright.
Actually not: the copyright claim would be particularly egregious in, e.g. the UK, because if the publisher has not even acknowledged or attributed contributions to Claude, then Claude's copyright is infringed, and the publisher could be liable for fraud on top of that.
And to address your second point: if the publisher claims 100% ownership, authorship and copyright, then who can even determine the amount of LLM-authored code? Something must come out in depositions or the courtroom about how much Claude committed, and how much was by humans, because in this case in this thread, the publisher has claimed 100% human authorship.
> Actually not: the copyright claim would be particularly egregious in, e.g. the UK, because if the publisher has not even acknowledged or attributed contributions to Claude, then Claude's copyright is infringed
I think you are wrong there. Claude cannot hold a copyright. Anthropic might but I think that is a misinterpretation of the law. The government thinks the person who prompted holds the copyright: "In the case of a general purpose AI which generates output in response to a user prompt, the “author” will usually be the person who inputted the prompt." :
Are you saying that this is wrong and the developer of the LLM holds the copyright on all its output?
> And to address your second point: if the publisher claims 100% ownership, authorship and copyright, then who can even determine the amount of LLM-authored code?
Does it matter? Either the human author/prompter holds the copyright, or no one does. They just need to show that they have made a sufficient contribution to hold the copyright. Commits and prompt history would be a good start.
https://news.ycombinator.com/item?id=49203613