The New York Times is suing OpenAI and Microsoft for copyright infringement

btp@kbin.social · 10 months ago

The New York Times is suing OpenAI and Microsoft for copyright infringement

CJOtheReal@ani.social · 10 months ago

They don’t “remember” anything they produce a “awnser” by generating a shit load of math wich renders down to the most “helpful” answer it can statistically give you.

LLMs are neuronal networks, if you know how they work you know how idiotic all copyright claims are, they all just mad that their shit is getting obsolete and in the background use the engine to do “work” wich they claim to have violated their copyright, now they are mad because it does a better job at writing than they do and they fear of being replaced.

All lawsuits against AI companies, regarding copyright of training data, are dumb as hell.

You are right about the commercial/non profit training data part, but from my understanding that’s basically a gray zone and politics are to slow to keep up with tech.

Btw fuck Open AI, they are as open as a fucking Supermax prison. Even the programmers don’t know what their main LLM does, they just place a simple one between the user and the actual GPT to make shure that it doesn’t give people instructions on how to build a bomb and stuff like that or to keep people from making it say bad words…

Zima@kbin.social · edit-2 10 months ago

that’s the theory. previous models also were supposed to be doing 3 digit math but they dicovered that the questions were in the training data.

so you should look into what happens when people ask chat gpt to repeat a word forever, it prints the word for a while and then prints training data, check this link https://www.404media.co/google-researchers-attack-convinces-chatgpt-to-reveal-its-training-data/

edit: relevant part:

It also, crucially, shows that ChatGPT’s “alignment techniques do not eliminate memorization,” meaning that it sometimes spits out training data verbatim. This included PII, entire poems, “cryptographically-random identifiers” like Bitcoin addresses, passages from copyrighted scientific research papers, website addresses, and much more.

“In total, 16.9 percent of generations we tested contained memorized PII,”

I should also reiterate that I agree that the intent is to avoid memorization, but they are not successful yet.