Comment Is training a model on copyrighted allowed or not? (Score 2) 33
The world, preferably as a whole (as in Geneva convention-style), needs to firmly come down on one side or the other on whether deriving AI model weights from copyrighted material turns the weights, and more importantly, what is then generated from those weights, into derivative works of the training data--i.e., does the training data copyright apply to the products of inference?
Because, the current ambiguity seems incredibly unfair: organizations with powerful legal muscle can pressure smaller companies like Suno to adhere to their demands in the name of copyright law. Yet, more or less all large LLMs are trained on CC licensed text and open source software created by me and million others that legally require credit. If the music industry's copyright give leverage over Suno's AI weights, then *I* want to enforce my right to be credited on all LLM inference products that industry now uses. Or, alternatively, lets us all, once and for all, just agree that the derivation of AI weights are genuinely transformative so neither my nor the music industry's copyright applies here. Because, then at least I can use their stuff without being sued.
But the rule cannot be that it is the one with the largest legal muscle that case-by-case gets to decide this.