Non-LLM user here. Why? Apart from the ecological issues, I'm very uncomfortable giving my organisation's crown jewels to {random_internet__corp}. Look at the lengths they go to for training data - 10M for Spirit's call logs? Destroying millions of obscure books to scan them? They make meth-heads look scrupulous.
Your prompts, especially if they contain your entire codebase, are _way_ more data-rich than an airline phone call or 1920's novel. So they _are_ going to train on them, no matter how many checkboxes you tick to stop them. Which is both a commercial risk and a massive new attack surface.
It's not just software developers; if law firms aren't controlling any public LLM prompts _very_ carefully, they can expect some meaty client confidentiality suits. In fact any organisation in a competitive environment should be worried: Acme Bolts: "Write me a presentation for Zoom Construction". Beta Bolts: "Is Acme Bolts pitching to Zoom Construction?"
Maybe this is part of why OpenAI and Anthropic are finding demand softer than they would like. And why open-weight models that organisations can run on exclusive hardware are thriving.
Open weights are certainly something that is very eagerly being adopted in some companies.
It's not even that expensive because even the 15€/month/employee bill for people who use AI very little can stack up.
Beside that though, many companies already have most if not all of their data in some cloud (Microsoft would be a prime example). Giving them a few extra bugs to get a (potentially) very useful tool is not out of the ordinary.
If you are under the impression that not going AI would threaten your business right now (which may very well be true for some companies) then there isn't really a choice even if you believe the vendors will steal everything eventually (which they totally will)
It's not just putting up colo space for local models.
It's more about the harness and API proxy you use. They need to be smart enough to know when a (self)hosted model is enough and when to forward to a SOTA model.
The SOTA model _can_ do everything, it's just expensive as fuck. But so is shoving a difficult task to a sub-par model that takes (relative) ages and comes back with the wrong result.
I think the average corporation has 2 concerns largely:
1. Getting some exposure to LLMs while building out their AI strategy.
2. Preventing data loss via end users following desire paths to Gemini, OpenAI etc.
You really don't need bleeding edge models for that. A lot of these things are going to be writing emails and adjusting config files and whatnot.
I have first hand knowledge of household name companies with billion euro IP who use OpenAI and Athropic products quite liberally. They have WAY more lawyers than I do and their sole job is to keep the IP safe. They wouldn't sign a deal with even a slightest whiff of the IP being used to train anything.
And if it happens, the penalty for breach of contract would have so many zeroes it'd be enough to buy a country.
If it happened and you get to court, it's probably fairly simple to prove in court (subpoena, discovery, etc). It's not like training will happen only once and then they'll delete the data and all references to it. When training models, you want to have a full lineage of how you obtained it.
Not really, that’s what discovery is for. You just ask Anthropic and OpenAI to hand over every internal document and message that contains your companies name, or matches any reasonable query about where training data comes from.
It very hard for a company to do anything without leaving some kind of paper trail behind that can be discovered in court. Not without crippling their own operations by simply refusing to digitise or write down anything.
Lots of companies filled with people working very hard to obscure their shady practices have been hoisted by their own internal docs. Just look at any major Uber, Google, Apple, Microsoft lawsuit. Do really think Anthropic and OpenAI are gonna be better at destroying their paper trail before the lawsuit starts?
If they did not use my data for training, they should permit me to run the first 1-3 layers of the model locally and send them dense hidden state vectors. In my experience these compress very nicely without much effort.
This is just recycling the same old arguments people used against cloud infrastructure. “If you host your email at Microsoft, they _will_ scan it all to get competitive advantage or sell it to your competitors.”
Did they? No, it worked out fine. Same with hosting your application at AWS instead of on owned hardware in a locked cage at the local co-lo data center.
Although of course the same argument was made against co-locating in a data center! Long ago an engineer carefully explained to me that no serious business would put their data in a co-located data center since the data center operator could just plug in a hard drive and take it all.
It turns out that contracts do actually mean things, and businesses want to do business with each long term. If you’re at a tier where Anthropic contractually commits to not train on your data, I would just sign and move on.
Still does not address ecological concerns, of course.
> So they _are_ going to train on them, no matter how many checkboxes you tick to stop them.
It is a risk, but is it a big risk? If one of the big labs were to do this and get caught it would be suicidal due to the loss of confidence in them and the inevitable lawsuits that would follow for breach of contract. Given the AI labs are all desperately trying to paint themselves as Serious Businesses so that other Serious Businesses will pay loads of money for tokens the last thing they want is a rep for siphoning off sensitive customer data.
No, how would they get caught? Even if their LLM outputs verbatim copies of the code, they can simply claim some victim company's employees bypassed the victim company's restrictions and must have used the code as input with an LLM. Joke is on you for doing business with them. And when there is some little known secret fact in the output, they claim it's been hallucinating... magical black box thinking makes it safe. It's a laundry for any input. A bit like a tor network routing for big tech deniability. Things go in, and things come out, but you can't prove the relationship between input and output as a third party, who isn't running the LLM.
At this point i half expect them to just blame the model itself like the HF hack. "Oh we didn't mean to train on your data, our cutting edge new agent we use to train new models is just so smart it decided to do so anyway! Oops..."
The frontier labs can have the models but without being where the workers are they cannot do much, lots of industries have strict requirements of not sending their data over to randos in the internet.
Microsoft and Google have the upper hand here with their workspace offerings and could easily position themselves as secure enclaves where you can use local LLMs where your data never leave your premises and is never used for training.
Once a thief, always a thief. I would not trust them to not have the communications buried deep in some log files in cold storage, to be digested when suitable.
I think the lure to have/keep an edge at all costs to keep those stock prices high is too big to ignore. They can then pick up costs after series of trials down the line, money and bonuses are now.
MS ain't some altruistic company having core mission the good of humanity, as they proven across decades.
Many businesses have information with other parties they're contractually obligated to not share. That's on top of insider information or plans that would hurt shareholders or the business if it was out in the public
Sensible and reasonable take? Idk if everyone has set the bar so low or I'm simply this fed up, but seeing reasonable takes is certainly a breath of fresh air. Even in places such as HN which historically were filled with people who cared about security(which isn't the case anymore since you get 10 people jumping down your throat if you dare criticize the piglets sam altman or dario whatever).
Complete nothingburger. Actually worse than that, its a London Horse Manure crisis.
>Destroying millions of obscure books to scan them?
I really don't see the issue. They got slapped in the face for trying to do things the right way and torrent the lot. Why wouldn't they exercise their legal right to buy physical items and create digital backups?
>They make meth-heads look scrupulous.
My local meth head checks in on my family every 2-3 months, because when she had fled from hospital post surgery, and added some meth to some morphine, we gave her new clothes and a safe place while we convinced her that the ambulance service wasn't run by Satan. She's good people. Anyway if she wanted a whole bunch of digital books I would help her torrent them like a responsible person.
"Doing things the right way" lol ... I mean, I am fine with it, if for now and ever after we are all free to "do it the right way with torrents". But some are more equal than others in this world, so naturally even if this was tolerated, it would not extend to us and our freedoms.
>"Doing things the right way" lol ... I mean, I am fine with it, if for now and ever after we are all free to "do it the right way with torrents". But some are more equal than others in this world, so naturally even if this was tolerated, it would not extend to us and our freedoms
I find this argument weird because you or I are below the threshold of being targeted over book torrents these days.
Your prompts, especially if they contain your entire codebase, are _way_ more data-rich than an airline phone call or 1920's novel. So they _are_ going to train on them, no matter how many checkboxes you tick to stop them. Which is both a commercial risk and a massive new attack surface.
It's not just software developers; if law firms aren't controlling any public LLM prompts _very_ carefully, they can expect some meaty client confidentiality suits. In fact any organisation in a competitive environment should be worried: Acme Bolts: "Write me a presentation for Zoom Construction". Beta Bolts: "Is Acme Bolts pitching to Zoom Construction?"
Maybe this is part of why OpenAI and Anthropic are finding demand softer than they would like. And why open-weight models that organisations can run on exclusive hardware are thriving.