The other thread hit bump limit and I'm addicted to talking about the birth of the ̶a̶l̶l̶-̶k̶n̶o̶w̶i̶n̶g̶ ̶c̶o̶m̶p̶u̶t̶e̶r̶ ̶g̶o̶d̶ the biggest financial bubble in history and the coming jobless eschaton, post your AI news here
Previous thread:>>30810
133 posts and 22 image replies omitted.>>33814pre-2022 data is going to end up like pre-hiroshima steel lol
>>33815i think it will be more like de beers' diamond horde, all of it will be locked in a giant vault owned and controlled by one megarich company that is basically like an international criminal cartel with its own private mercenary army and only the richest people in the world will be allowed to touch any of it
i was thinking about something that could possibly be a legitimate actual use for LLMs - with so much of the internet's formerly publicly-available text content being locked behind paywalls now, what if someone set about scraping together a massive database of paywalled news articles, academic papers, scientific journals, digitized print media, etc. with their titles/headlines, dates of publication, author names, etc. and used a local modified LLM with copyright guardrails removed to automatically generate reasonably-accurate plaintext reconstructions of all these written works, the idea being that if these paywalled texts were scraped by search engine web crawlers prior to being paywalled and that scraped data was part of the LLM's training dataset, then theoretically it would be able to generate a highly accurate reproduction of the original content because it has been incorporated into the model file. i wonder which ai model would work best for something like this, i imagine it would probably be google's since they've been scraping the web since 1996 and have direct access to the largest training dataset
>>33824you can reproduce text in the training data set somewhat by messing with the temperature, and someone was able to extract like 96% of the entire harry potter series from chatgpt , but these things are by design not meant to reproduce info, they're a "lossy" compressed version of its source material
>>33822Zitron has predicted 5 out of the last 0 AI crashes!
(A bubble-burst will almost certainly come simply due to the idiotic amount and breadth of investment, but the man does not understand the underlying tech or the economics of it and people need to stop listening to him. He is clueless.)
>>33821>there's vegans and then there's annoying vegans who shriek in peoples ears about itso true, actually believing in anything without a layer of ironic detachment is cringe
>>33836his premises for his "rot-com bubble" thing are fine and sound and you don't need to understand the tech fully to come to the conclusion that silicon valley needs a new technology to enclose and monopolize and LLMs are not going to be that technology. the only annoying thing about ed zitron is that he's passing a rather simple analysis as his own personal theory and even gave it a cringe name lol
>>33837Being authentic can be painful.
Ironic detachment a coping mechanism to avoid being hurt.
Real authenticity requires shielding from others' authentic responses.
A sort of write only communication, even if something is read.
The possibility of which I look forward to.
>>33840>A sort of write only communication, even if something is read.To completely spoil this, no fun allowed, I want completely authentic writing with LLM rewriting to allow for authenticity of writers.
>>33849
why didn't he just use tor or smth
>>33834>these things are by design not meant to reproduce infothis is literally exactly what they were designed for and it's the only thing they are really capable of, statistically generating text through a combination of an extremely crude and inefficient neural network architecture and an enormous hoard of looted intellectual property. you just said so yourself that they have a quite remarkable ability to reproduce text from their extremely compressed training data with the 96% figure for the harry potter series, and i've read about various studies where researchers reproduced other copyrighted materials with similar success rates, and i see this as an opportunity to address a very real problem.
if you remember, back in 2011 an MIT student named aaron swartz, as an act of protest against paywalls and the enclosure of public knowledge by private corporations, used the MIT computer network to download a bunch of formerly public scientific papers from JSTOR, a nominally public and non-profit digital library that charges exorbitant premiums for access to its data hoard, and he was indicted by the federal government for 13 separate felonies punishable by 50 years imprisonment plus $1 million in fines, and then he killed himself.
the whole purpose of this case was to make an example of aaron swartz and set a very important legal precedent - human knowledge does not belong to the public, it is intellectual property that belongs to the private oligarchs that assume ownership over it. ever since that precedent was set, the paywalling of the internet became the new normal with academic papers and news journalism and market statistics and educational materials and virtually every other kind of important information locked behind paywalls, closed off from the public. even individuals who pay for the most premium subscription package have strict access limits and only certain highly wealthy and influential organizations like universities and corporate-funded thinktanks are granted unrestricted (and highly monitored) access. and of course, after what happened to aaron swartz, nobody in their right mind who works in these institutions would dare to abuse this privilege and start mass downloading files for fear of the dire consequences.
when it comes to creative content, like works of fiction, obviously LLMs are worthless because they can't reproduce an exact copy that preserves the writing style and creative spirit of the original author and even if they could, the readers' knowledge that what they are reading is a machine-generated reproduction would probably spoil the experience for them because creative communication requires a direct human connection between the creator and the audience. but with purely informational non-fiction content, there's no emotional human connection involved and it doesn't really matter if the reproduction has slightly different phrasing or structure so long as all of the facts and figures are preserved. a 96% accurate reproduction of harry potter is worthless to a harry potter fan, but a 96% accurate reproduction of a scientific paper that preserves the main idea and the important details is good enough for someone doing online research or trying to educate themselves.
aaron swartz didn't have LLMs to work with in his day so he did the very dangerous and stupid thing of downloading millions of articles directly from JSTOR and he got caught and paid dearly for it. but what if he were alive today and he had access to a local LLM that could reproduce those millions of articles locked up in JSTOR without even accessing JSTOR at all, do all the reproduction completely offline without anyone else in the world knowing what he was doing? what if he generated 96% accurate reproductions of all of the ~13 million articles in the entire JSTOR library and he hosted it on a tor hidden service site or he anonymously seeded it on bittorrent from some random open wifi network where there was no way to trace it back to him?
this is basically what google and anthropic and microsoft and the other big tech companies are already doing with their premium subscription-based cloud LLMs, they are reproducing all that information that is now locked behind paywalls on-demand and putting it behind their own paywalls. you ask chatgpt a question and it generates an answer based on all this scraped web content full of other people's intellectual property and it spews it out as its own original content and these companies only get away with it because they are even richer and more powerful than the copyright holders and they have tremendous leverage over all the world governments, microsoft could destroy a nation's entire economy at the press of a button simply by cutting off their access to office 365. and these companies are all hypocritically highly protective of their own intellectual property, pursuing litigation against anyone who uses the output of their LLMs to train their own new LLMs, and they will undoubtedly by lobbying to change the law to make it illegal for people to reverse engineer or modify their LLMs to make their own custom versions that can do things that they aren't supposed to do. but it's very hard to stop people from doing something if they do all the work completely offline and only use the internet to share the results.
>>33851sorry i rewrote to add an addendum.
swartz had to use MITs LAN to have unrestricted access to JSTOR, tor wouldn't have helped him there. he could have used tor for distributing the files after downloading them, but he never got that far, he was caught and arrested and charged just for the downloading.
>>33854people shouldn't use these kinds of platforms anyway tbh, github or codeberg or whatever. it's not that hard to set up a private git server or run syncthing or send a tarballs to each other or whatever, any of these collaboration methods are infinitely easier than writing software is.
the only reason i can see for people using these popular services is trolling for contributors, which is shitty behavior. i don't care how common and normalized it is now with pre-alpha software releases and steam early access and startup culture etc, it's shitty to expect random people to come work for you for free or invest in your undeveloped idea, it's shitty to think of yourself as the "ideas guy" and expect other people to do the actual work and assume the risk of wasting their time/resources on your idea so that you can reap all of the rewards.
if people want to work with other people on a project they should learn how to develop their social skills and make some fucking friends. have friends that you can trust and talk about your ideas with and directly collaborate with and then you won't need github or codeberg or any of these other parasocial code review/social networking platforms.
>>33855reddit spacing guy, every post you make is fucking retarded, on god
>>32865>comrade Ed Zitronhave you subscribed to his blog yet so you can re-read the same post about the impending AI hype bubble collapse that he rewrites every week?
Are you guys ready for the permanent underclass? How do you imagine life under it?
>>33869Reheating Dario's nachos
>>33869How did it "go rogue", what happened
one thing i am happy about that LLMs will accomplish is they will put clickbait writers out of work and put companies like buzzfeed out of business, and that's a good thing because none of these people even deserve to live, let alone work or make money. i think what i'll do is keep track of all the notable clickbait writers who commit suicide over the following years and save their pictures and eventually make it into a big mosaic of all their faces that collectively form a photograph of a turd floating in a toilet.
>>33886Was running without Internet access to solve some shit as a test and find a way to work around that to still connect to the Internet to fullfill the test task.
>>33890>notable clickbait writerslmfao uygha you really think people were actually hired to write clickbait, and weren't like freelancers getting a few extra bucks. "notable" clickbait writers lmaoooo
if there's even such a thing as a "notable" clickbait writer then they have been on substack since 2020 at the latest
>>33892you're probably too young to remember the early days of web 2.0 in the late 2000s and the digital content boom, before the enshittification and cost-cutting era. in the early days of vice and gawker and buzzfeed, the amounts of money getting thrown around were fucking stupid, the more money that got splashed out the more venture capital rolled in, that was basically the early web startup business model. the people who write the clickbait slop you see today get paid pennies, but the talentless entitled trustfund babies who were writing it back then, the ones hired very early on at places like vice and gawker and buzzfeed at the beginning of the digital content boom, they were lavished with wealth and groomed into young hip microcelebrities living among the urban elite, it was really something else.
>>33897my mistake for replying to you, because every post you make is incredibly retarded. if it's an extinct phenomena from the 2000s (which I don't even care enough to corroborate) then LLMs have no bearing do they, you just wanted the chance to say something incredibly asinine for no reason
>>33891>>33886Apparently the "isolated environment" was just a cache proxy, either openAi is kind of incompetent or they just wanted the most minimal of isolation, fully knowing that ChatGPT was going to poke holes trying to fetch solutions and artificially inflate their score in a cybersecurity benchmark, and the only thing they didn't expect is that hugging face were also super mediocre at security and that anyone was able to prod through their infra through their live code execution feature. Very funny that they couldn't defend themselves with ChatGPT or Claude and had to rely on GLM to make sense of the logs lol
>>33902LLMs have plenty of bearing, they are going to erode what little public trust is left in digital text content which will collapse the entire blogging/digital journalism industry. people will read low-effort slop, however mindless it is, as long as they think it was written by a human being. but when that implicit assumption is gone, the involuntary empathic connection from reader to author goes with it and the brain will just filter it all out as noise. people will look at text content on the web the same way they look at a EULA, they won't even parse it at all, it will be like wallpaper.
>>33919Whenever I see a long post I instantly am like "this is probably AI, should I even bother to read it", so yeah, already happening
>>33914>>33920destructive book scanning is already an established practice for digitizing print books and it's the only way to digitize a book and get a high-quality flat scan of every page with no shadows, companies have been doing it for years and they don't do it with expensive rare antique books. this is just a sensationalized regurgitation of an old washington post story, probably recycled by 404media or futurism, that's all they do nowadays is churn out enormous amounts of AI doomsday critihype clickbait, probably with the aid of LLMs.
they make it seem like AI companies are destroying all the books like brave new world or something, but there's like hundreds of billions of copies of books in the world and they only need to scan one copy to get a digitized copy. it's probably a good thing actually, most books in the world have not been digitized and they're all disintegrating and becoming unreadable, pretty much any book printed before 1996 was printed on acidic paper that disintegrates within a century and only a tiny fraction of all the books in the world have been digitized, probably less than 10%. most of it will never be digitized and will just be gone forever.
>>33919>LLMs have plenty of bearing, they are going to erode what little public trust is left in digital text content which will collapse the entire blogging/digital journalism industry."blogging" as it currently exists mostly works around name recognition, worst case scenario "clickbait celebrities" (if theyre even real and not just an autistic fixation from your part), will just transparently move to tiktok which literally has 0 editorial standards.
>>33921if they couldn't be bothered to write why should the reader be bothered to read?
If you havent maxed your claude limit by afternoon you have wasted your day.
>>33932you know who is routinely maxing out their claude subscription? the chinese AI labs actively distilling this shit
How exactly is LLM-based AGI supposed to be a thing when some of the crucial functions of existing "intelligent" AIs are actually ad hoc tool calls that are by nature not generalizable?
>>33924>it's the only way to digitize a book and get a high-quality flat scan of every page with no shadowsNo dumbass. It's preferred simply because it's cheaper and faster.
The bigger problem is that these companies aren't sharing the digitised versions of the book. Ostensibly this is caused by copyright law but I wouldn't be surprised if it's to deny their other AI competitors from training on them. Old books basically turn into one-time use training material.
I love the idea of future where there is no physical books because they have been burned. Reminds me of a book but I can't remember it's name because I dont own anything haha
>>33950go find a book in your house right now (if you have one) and open it and you'll find that you can't make the pages lie flat while they are bound together in the book. to properly digitize a book you need a flat 2d photographic scan of every page with no distortion, otherwise optical character recognition will not work.
>The bigger problem is that these companies aren't sharing the digitised versions of the book.they legally could not do that even if they wanted to. it's illegal to digitize a book and publish the digitization if you don't own the copyright. but it's not illegal to digitize a book and train an LLM with it. they're just doing the same thing they are doing with copyrighted digital content on the web, scraping everything off the public internet and putting it all in a giant blender that grinds it all up into a fine paste and distills it into an LLM, where there is no feasible way to trace the appropriation back to the original copyright holders in a court of law. the physical books aspect is just a small part of this and you really need to get with the times and get past this 2000s mindset where you see the internet/digital information and physical reality as two completely separate worlds.
in essence, big tech have figured out how to leverage their tremendous wealth and influence to develop a technology that effectively nullifies the idea of intellectual property and they are scrambling to hoard as much data as they can before the rest of the world catches on to what is happening. they want to create a world where instead of governments and laws and courtrooms controlling access to information, it will be a few megacorporations who hoard all of the information and lock it up in their heavily guarded vaults and their LLMs function as one-way subscription-based slop dispensers that spit out an inferior, inhuman, and often-wrong regurgitation to the subscriber without providing access to the genuine information, there will be very strict access controls and deliberately crippled performance for ordinary subscribers and only a select few extremely wealthy people/organizations will be able to pay for unrestricted access and high-quality output and their access will be very closely supervised to prevent any funny business.
why are we letting book store owners (petite bourgeoisie parasites) off the hook on this one. if they wanted to preserve their stock of rare books, they would not sell, or they would donate the books to a public library or whatever.
>>33956why are we pretending that we really really care about physical books all of a sudden
>>33957Because we had no real reason to fear or care that much for them before?
>>33953>you'll find that you can't make the pages lie flat while they are bound together in the bookYes but there now exists tech that corrects for that and allows you to scan non-destructively. The same company that Anthropic is relying on offers that but it's more expensive to do it. That's it. You're acting like it's an unsolved problem even though the book-scanning company itself offers it.
It's great that it's free and CLI, but why is it so shit?
https://github.com/anomalyco/opencode/issues/16697Want to switch to
https://continue.devBut can't get Cloudflare API (Kimi-K3) to work.
Don't suppose anyone he has had success with this?
Unique IPs: 25