overview for Eccitaze

2 authors say OpenAI 'ingested' their books to train ChatGPT. Now they're suing, and a 'wave' of similar court cases may follow. in c/[email protected]

[–] [email protected] 0 points 2 years ago (1 children)

Is comprehension necessary for breaking copyright infringement? Is it really about a creator being able to be logical or to extend concepts?

I think we have a definition problem with exactly what the issue is. This may be a little too philosophical but what part of you isn’t processing your historical experiences and generating derivative works? When I saw “dog” the thing that pops into your head is an amalgamation of your past experiences and visuals of dogs. Is the only difference between you and a computer the fact that you had experiences with non created works while the AI is explicitly fed created content?

That's part of it, yes, but nowhere near the whole issue.

I think someone else summarized my issue with AI elsewhere in this thread--AI as it currently stands is fundamentally plagiaristic, because it cannot be anything more than the average of its inputs, and cannot be greater than the sum of its inputs. If you ask ChatGPT to summarize the plot of The Matrix and write a brief analysis of the themes and its opinions, ChatGPT doesn't watch the movie, do its own analysis, and give you its own summary; instead, it will pull up the part of the database it was fed into by its learning model that relates to "The Matrix," "movie summaries," "movie analysis," find what parts of its training dataset matches up to the prompt--likely an article written by Roger Ebert, maybe some scholarly articles, maybe some metacritic reviews--and spit out a response that combines those parts together into something that sounds relatively coherent.

Another issue, in my opinion, is that ChatGPT can't take general concepts and extend them further. To go back to the movie summary example, if you asked a regular layperson human to analyze the themes in The Matrix, they would likely focus on the cool gun battles and neat special effects. If you had that same layperson attend a four-year college and receive a bachelor's in media studies, then asked them to do the exact same analysis of The Matrix, their answer would be drastically different, even if their entire degree did not discuss The Matrix even once. This is because that layperson is (or at least should be) capable of taking generalized concepts and applying them to specific scenarios--in other words, a layperson can take the media analysis concepts they learned while earning that four-year degree, and apply them to a specific thing, even if those concepts weren't explicitly applied to that thing. AI, as it currently stands, is incapable of this. As another example, let's say a brand-new computing language came out tomorrow that was entirely unrelated to any currently existing computing languages. AI would be nigh-useless at analyzing and helping produce new code for that language--even if it were dead simple to use and understand--until enough humans published code samples that could be fed into the AI's training model.

2 authors say OpenAI 'ingested' their books to train ChatGPT. Now they're suing, and a 'wave' of similar court cases may follow. in c/[email protected]

[–] [email protected] -2 points 2 years ago (2 children)

You realize LLMs are designed not to self improve by design right? It’s totally possible and has been tried - It’s just that they usually don’t end up very well once they do.

Tay is yet another example of AI lacking comprehension and intelligence; it produced racist and antisemitic content because it had no comprehension of ethics or morality, and so it just responded to the input given to it. It's a display of "intelligence" on the same level as a slime mold seeking out the biggest nearby source of food--the input Tay received was largely racist/antisemitic, so its output became racist/antisemitic.

And LLMs do learn new things, they’re just called new models. Because it takes time and resources to retrain LLMs with new information in mind. It’s up to the human guiding the AI to guide it towards something that isn’t copyright infringement.

And the way that humans do that is by not using copyrighted material for its training dataset. Using copyrighted material to produce an AI model is infringing on the rights of the people who created the material, the vast majority of whom are small-time authors and artists and open-source projects composed of individuals contributing their time and effort to said projects). Full stop.

Also, you say “right” and “probable” are without difference, yet once again bring something into the conversation which can only be “right”. Code. You cannot create code that is incorrect or it will not work. Text and creative works cannot be wrong. They can only be judged by opinions, not by rule books which say “it works” or “it doesn’t”.

Then why does ChatGPT invent Powershell cmdlets out of whole cloth that don't exist yet accomplish the exact precise task that the prompter asked it to do?

The last line is just a bit strange honestly. The biggest users of AI are creative minds, and it’s why it’s important that AI models remain open source so all creative minds can use them.

The biggest users of AI are techbros who think that spending half an hour crafting a prompt to get stable diffusion to spit out the right blend of artists' labor are anywhere near equivalent to the literal collective millions of man hours spent by artists honing their skill in order to produce the content that AI companies took without consent or attribution and ran through a woodchipper. Oh, and corporations trying to use AI to replace artists, writers, call center employees, tech support agents...

Frankly, I'm absolutely flabbergasted that the popular sentiment on Lemmy seems to be so heavily in favor of defending large corporations taking data produced en masse by individuals without even so much as the most cursory of attribution (to say nothing of consent or compensation) and using it for the companies' personal profit. It's no different morally or ethically than Meta hoovering all of our personal data and reselling it to advertisers.

2 authors say OpenAI 'ingested' their books to train ChatGPT. Now they're suing, and a 'wave' of similar court cases may follow. in c/[email protected]

[–] [email protected] -1 points 2 years ago (2 children)

Undertale was allowed to exist because none of the elements it took inspiration from were eligible for copyright protection. Everything that could have qualified for copyright protection--the dialogue, plot, graphical assets, music, source code--were either manually reproduced directly by Toby Fox and Temmie Chang, or used under permissive licenses that allowed reproduction (e.g. the GameMaker Studio engine). Meanwhile, the vast majority of content OpenAI used to feed its AI models were not produced by OpenAI directly, nor were they obtained under permissive license.

So... thanks for proving my point?

2 authors say OpenAI 'ingested' their books to train ChatGPT. Now they're suing, and a 'wave' of similar court cases may follow. in c/[email protected]

[–] [email protected] 0 points 2 years ago (4 children)

"right" and "probable" text are distinctions without difference. The simple fact is that an AI is incapable of handling anything outside its learning dataset. If you ask an AI to talk like a pirate, and it hasn't had any pirate speak fed to it by a human via its training dataset, it will utterly fail. If I ask an AI to produce a Powershell script, and it hasn't had code fed to it by a human via its training dataset, it will fail utterly. An AI cannot proactively buy a copy of Learn Powershell In a Month of Lunches and teach itself how to use Powershell. That fundamental shortcoming--the inability to self-improve, to proactively teach itself and apply that new knowledge to existing concepts--is a crucial, necessary element of transformative effort required to produce a derivative work (or fair use).

When that happens, maybe I'll buy that AI is anything more than the single biggest copyright infringement scheme the world has ever seen. Until then, though, I will wholeheartedly support the efforts of creative minds to defend their intellectual property rights against this act of blatant theft by tech companies profiting off their work.

2 authors say OpenAI 'ingested' their books to train ChatGPT. Now they're suing, and a 'wave' of similar court cases may follow. in c/[email protected]

[–] [email protected] 1 points 2 years ago* (last edited 2 years ago) (10 children)

Again, that's not comprehension, that's mixing in yet more data that was put into the model. If you ask an AI to do something that is outside of the dataset it was trained on, it will massively miss the mark. At best, it will produce something that is close to what you asked, but not quite right. It's why an AI model that could beat the world's best Go players was beaten by a simple strategy that even amateur Go players could catch and defeat--the AI never came across that strategy while it was training against itself, so it had no idea what was going on.

And fair use isn't the bulletproof defense you think it is. Countless fan games have been shut down over the decades, most of them far more transformative than my hypothetical example, such as AM2R. You bet your ass that if I tried to profit off of that hypothetical crossover roguelike, using sprites, models, and textures directly ripped from their respective games, it would be shut down immediately.

EDIT: I also want to address the assertion that AI isn't trained to recreate existing works; in my view, that's wholly irrelevant. If I made a program that took all the Harry Potter books, ran each word through a thesaurus, and sold it for profit, that would still be infringing, even if no meaningful words were identical to the original source material. Granted, if I curated the output and made a few of the more humorous excerpts available for free through a Mastodon or Lemmy post, that would likely qualify as fair use. However, that would be because a human mind is parsing the output and filtering out the 99% of meaningless gibberish that a thesaurus-ized Harry Potter would result in.

The only human input to an AI that gave consent to being part of its output is the miniscule input of the prompt given to it by the human, which does not meet the minimis effort required for copyright protection under law. The rest of the input--the countless terabytes of data scraped from the internet and fed into the AI's training model--was all taken without the author's consent, and their contribution vastly outweighs that of the prompt author and OpenAI's own transformative efforts via the LLM.

2 authors say OpenAI 'ingested' their books to train ChatGPT. Now they're suing, and a 'wave' of similar court cases may follow. in c/[email protected]

[–] [email protected] 2 points 2 years ago (15 children)

The problem with AI as it currently stands is that it has no actual comprehension of the prompt, or ability to make leaps of logic, nor does it have the ability to extend and build upon existing work to legitimately transform it, except by using other works already fed into its model. All it can do is blend a bunch of shit together to make something that meets a set of criteria. There's little actual fundamental difference between what ChatGPT does and what a procedurally generated game like most roguelikes do--the only real difference is that ChatGPT uses a prompt while a roguelike uses a RNG seed. In both cases, though, the resulting product is limited solely to the assets available to it, and if I made a roguelike that used assets ripped straight from Mario, Zelda, Mass Effect, Crash Bandicoot, Resident Evil, and Undertale, I'd be slapped with a cease and desist fast enough to make my head spin.

The fact that OpenAI stole content from everybody in order to make its model doesn't make it less infringing.

Outrage as Republican says 1921 Tulsa massacre not motivated by race in c/[email protected]

[–] [email protected] 3 points 2 years ago

Well, so much for that........

Psycho ex-partner in c/[email protected]

[–] [email protected] 3 points 2 years ago

Maybe consider digitizing that cassette, or at least listening to it to make sure it's still usable (assuming you can stomach it). Cassette mediums degrade over time and it's quite possible that microcassete could be reaching the end of its usable life: https://en.m.wikipedia.org/wiki/Preservation_of_magnetic_audiotape