I think there is room to argue though, that distillation might violate copyright while training on large volumes of large material does not.
At that point you are modeling a model. I think you could argue that it is a bit like copying a novel where you hand type it; adding lots of typographical errors a long the way; get laze and leave out a chapter or two that you don't feel moves the story a long much, then slap a new title on a call it your own.
I don't think any court anywhere would say did not violate the original authors copyright. This is why we have courts someone has to effectively make an analog for events that happen that reflect existing laws and convince a jury or at least a judge that analog is correct. As much as everyone wants to believe that is some kind of perfect deterministic process, but not really. If it was we would not have appellate courts, and the SCOTUS would never reverse itself.
Honestly about the only compelling argument for how these cases will go; IMHO is that "Si Valley Bros have a lot of cash to spend on lawyers, more than most of the plaintiffs or defendants they might face; that is usually favorable."
I do honestly believe courts usually get it right, They are run by well educated people who are genuinely gifted at absorbing information quickly in most cases. I think we can see this in complex cases where we have expertise; think about the Google implementing Java APIs case for example or the SCO's nonsense copyright claims on Linux. Despite enormous effort and resource commitment to cloud the issues, the courts saw thru the BS cloud. What is important there though is those cases were broadly speaking not new territory, the arguments and key decision points of law were well known.
AI models likely represent a need to invent new "legal tests" and frankly I don't see how anyone can be certain what the outcomes will be.