Not Navier-Stokes, that was a couple days ago, old news. It’s one of the less famous problems they posted around the same time.
The cited Mathstodon post is here: https://mathstodon.xyz/@andreasthom/117240535270608201
While this general story has its shape. The line at the end here…
AI is stealing human discovery.
Is striking in just how obviously true it’s been the whole time. It’s the business model, plain and simple.
I hope this story highlights this better for some. But also, it’s frustrating that any of this is new or surprising to anyone.
It’s one thing to notice OpenAI is stealing public content, but in this case it’s suspected of using info that people prompted ChatGPT with privately.
Do people consider conversations with LLMs private? I don’t think there is actual guarantee which protects the conversations. …are people that naive?
Yep.
The era of Google and big social media did real harm IMO. I suspect we used to have more scepticism around this sort of thing (dunno though). But getting addicted to things sure as hell is a good way to get you to ignore their problems.
The era of desktop computing really is gross in so many ways once you take a step back.
Huh? Desktop computing is just fine as long as you avoid Windows. It’s the era of mobile computing that brought on all the problems.
Desktop computing is being killed off. So is self-hosting.
LLMs being used to wow greedy yuppie capitalists into buying in so they can replace all workers and get huge returns next quarter for eliminating the companies most expensive cost center: people. “The buildout must be intense! We can’t get beaten by the Chinese!! Quick, order the entire output of the world’s computer components”
Now suddenly you can’t buy them. Do you think that’s an accident? Do you think they just naively did that? These guys who built these startups and these tools? These guys who figured out how to extract billions of dollars in personal wealth from the US economy - you think they didn’t know they were taking computers away from us?
Desktop computing is fucking dead my friend, unless they fail. But it’s looking less and less like anyone will or even can stop them.
I hear you, but I’ve started to wonder about what things are wrong intrinsically in digital tech, or at least that are hard to fight.
With a desktop computer that’s supposed to help you with everyday things, you’ve got to have an OS and software. I think that naturally leads to vendor lock in, but not just for replacement parts as it would be with a car but your personal information and artefacts.
I suspect it’s in the nature of building up a software stack that’s user friendly enough for mass consumption: there are just too many choices the developers need to make, and too many layers, and so too many opportunities a company has to make their software incompatible with someone else’s.
I suspect with desktop computing every day mental work was colonised by monopolies and monopolistic structures. Which then sent us down a problematic path which may very well strongly resemble cyber punk dystopias.
That big social media and then this AI surrender followed, is to me, not a coincidence at all.
There’s something funny about software that betrays its promise, a dangerous magic if you will.
I would expect that email sent through something like gmail is private. Even if there’s some possibility of interception under search warrants or stuff like that, I wouldn’t expect the email text to be used for training Google AI. I don’t use gmail (at least from the client side) anyway, but I think there was some noise about this a while back. Certainly I’d expect a paid LLM server to not train on the prompts. I don’t use those servers though, and haven’t checked the TOS. I might fool around with some open models sometime but have no interest in ChatGPT.
Gmail has been reading your mail and processing it for marketing data since day 1. This is made obvious by the fact that they sorted your mail into categories.
You didn’t really think that’s all they did with that data, did you?
Oh I didn’t know that! Thank you!
So the mathematicians were using ChatGPT in their work and that’s how they think it was scooped?
Thoigh to be fair, this doesn’t change my tune. That we aren’t all presuming every interaction will be part of future training data is bonkers to me.
That’s exactly the kind of data that turned LLMs into chat interfaces. And where else are AI companies going to get human generated text and feedback?
The extent to which the keys of the world have been handed over to a skynet system is fucking crazy!
stealing public content
How does one steal what is public?
This article is about stuff that was private, yes?
OpenAI is guilty of computer crime through massive unauthorized access to web servers to scrape their content, evading every kind of blocking attempt and crashing servers all over the place. This is the same thing Aaron Swartz was prosecuted for, but nothing seems to be happening to OpenAI. The content being public is irrelevant to this crime. You could have a public domain book in your bedroom, but if I break into your house and copy it, I’m a burglar even if I’m not a copyright infringer.
Meta (Facebook) is known to have done a huge copyright infringement by downloading a massive pirate library (libgen) to train its AI. That’s separate from the claim that using the texts for AI training is infringing in its own right (there are lawsuits about that going on, and it’s not a slam dunk issue imho). I remember this reported specifically about Meta but it would shock me if OpenAI and Anthropic didn’t do the same thing. So they have no business whining about model distillation.
Separately from that, yes, the stuff in the math prompts was supposed to be private but OpenAI apparently trained on them anyway. OpenAI is a multi-tasker and can do more than one bad thing at the same time.
Well, one obvious example is licensing. You can have things that are presented online with a terms of service or an actual license that excludes certain uses.
One specific example would be GitHub and the sites like it. There are plenty of public facing code bases with licensing that would prohibit forms of reuse, like for financial gain.
And yes, this appears to be what the researchers believed to be private conversations, but are likely completely owned by the AI companies in their terms of service.
Stealing essentially all of the info on the internet and attempting to close the door behind them so they can charge you metered access to what is essentially the library of Alexandria.
Let this be a warning to everyone that if you want to use AI assistance in secretive work, do so with local AI. If you must use a hosted AI, ensure your inference provider is not also the foundation model provider.
I know that this is unrelated, but I’m glad that nitter is back.
Awesome. Any known public instance working?
I mean, xcancel as linkedin the post :p
Maybe those mathematicians should stop using the proof-stealing slop machine then.
It’s hard, the frontier models are only available on the theft machines right now. IDK what will happen with that.



