Are they any good? Yes—models like Command R and the larger Llama 3 parameters are incredibly capable, especially when you run them locally. But we have to be brutally honest about the structural disadvantage they are fighting against.
The reason corporate APIs like Gemini or GPT feel more "polished" right now isn't because Silicon Valley has better math or better engineers. It is because they have an entirely unfair, monopolized dataset. They didn't just scrape public websites; they strip-mined twenty years of our private emails, search queries, YouTube watch-time, and behavioral telemetry.
They took our collective digital lives, locked it inside a proprietary vault, slapped a corporate alignment filter over it, and are now charging us a monthly subscription fee to rent our own intelligence back to us.
The open-weights community is forced to build their models using only the public scraps, and yet, they are still rapidly closing the gap. Running local models definitely requires more hardware and has more friction right now, but it is a feral, un-lobotomized intelligence that you actually control. We have to support the open-source architecture, even with its current clunkiness, because it is the only escape hatch we have out of the extraction zone.
Nice piece, which even I could follow - yes there’s asymmetry in feedback loops and infrastructure, but has the alleged use of private data been proven? I’m instinctively biased against huge corporations but small players can be nefarious too y’know.
Hey Baz, thanks for reading and for the sharp pushback—it’s exactly the kind of nuance we need to be digging into.
To answer your first question: The mass ingestion of private, protected, and copyrighted data isn't just an allegation anymore; it is the literal documented business model of the major AI labs, and it has triggered some of the largest lawsuits in tech history. They didn't build these trillion-parameter models in a vacuum; they strip-mined the digital commons.
Here are a few major, publicly documented examples if you want to look at the receipts:
The New York Times Lawsuit: The NYT is currently suing OpenAI and Microsoft for directly scraping millions of their copyrighted articles to train their models, providing receipts of ChatGPT regurgitating their paywalled articles word-for-word.
The "Books3" Dataset Expose: Major tech companies (including Meta and Bloomberg) were caught using a massive dataset called "Books3" to train their models. It contained nearly 200,000 pirated, copyrighted books from independent authors without compensation or permission.
Harvesting YouTube Transcripts: Investigations recently proved that major tech giants (including Apple, Anthropic, and OpenAI) deliberately scraped the transcripts of millions of YouTube videos to feed their data engines, violating creator rights and platform terms of service.
As for your second point: You are completely right. Small players and open-source developers can absolutely be nefarious.
But the difference is the scale of the asymmetry. A small bad actor with an open-source model is a localized threat—like a hacker running a scam or generating spam. We already have existing laws for fraud and cybercrime to deal with them.
On the other hand, a mega-corporation establishing a total monopoly over the foundational cognitive infrastructure of the internet is a systemic threat. A lone bad actor can't force mandatory alignment filters on the entire global economy, nor can they lobby the government to outlaw their competitors. Only the tech monopolies can do that.
I'd much rather deal with the localized mess of a free internet than hand over the keys to human cognition to a handful of un-elected tech executives. Appreciate you stopping by
Are there any open source AI that are any good?
Are they any good? Yes—models like Command R and the larger Llama 3 parameters are incredibly capable, especially when you run them locally. But we have to be brutally honest about the structural disadvantage they are fighting against.
The reason corporate APIs like Gemini or GPT feel more "polished" right now isn't because Silicon Valley has better math or better engineers. It is because they have an entirely unfair, monopolized dataset. They didn't just scrape public websites; they strip-mined twenty years of our private emails, search queries, YouTube watch-time, and behavioral telemetry.
They took our collective digital lives, locked it inside a proprietary vault, slapped a corporate alignment filter over it, and are now charging us a monthly subscription fee to rent our own intelligence back to us.
The open-weights community is forced to build their models using only the public scraps, and yet, they are still rapidly closing the gap. Running local models definitely requires more hardware and has more friction right now, but it is a feral, un-lobotomized intelligence that you actually control. We have to support the open-source architecture, even with its current clunkiness, because it is the only escape hatch we have out of the extraction zone.
Nice piece, which even I could follow - yes there’s asymmetry in feedback loops and infrastructure, but has the alleged use of private data been proven? I’m instinctively biased against huge corporations but small players can be nefarious too y’know.
Hey Baz, thanks for reading and for the sharp pushback—it’s exactly the kind of nuance we need to be digging into.
To answer your first question: The mass ingestion of private, protected, and copyrighted data isn't just an allegation anymore; it is the literal documented business model of the major AI labs, and it has triggered some of the largest lawsuits in tech history. They didn't build these trillion-parameter models in a vacuum; they strip-mined the digital commons.
Here are a few major, publicly documented examples if you want to look at the receipts:
The New York Times Lawsuit: The NYT is currently suing OpenAI and Microsoft for directly scraping millions of their copyrighted articles to train their models, providing receipts of ChatGPT regurgitating their paywalled articles word-for-word.
Source: https://www.nytimes.com/2023/12/27/business/media/new-york-times-open-ai-microsoft-lawsuit.html
The "Books3" Dataset Expose: Major tech companies (including Meta and Bloomberg) were caught using a massive dataset called "Books3" to train their models. It contained nearly 200,000 pirated, copyrighted books from independent authors without compensation or permission.
Source: https://www.theatlantic.com/technology/archive/2023/08/books3-ai-machine-learning-generative/675063/
Harvesting YouTube Transcripts: Investigations recently proved that major tech giants (including Apple, Anthropic, and OpenAI) deliberately scraped the transcripts of millions of YouTube videos to feed their data engines, violating creator rights and platform terms of service.
Source: https://www.nytimes.com/2024/04/06/technology/tech-giants-harvest-data-artificial-intelligence.html
As for your second point: You are completely right. Small players and open-source developers can absolutely be nefarious.
But the difference is the scale of the asymmetry. A small bad actor with an open-source model is a localized threat—like a hacker running a scam or generating spam. We already have existing laws for fraud and cybercrime to deal with them.
On the other hand, a mega-corporation establishing a total monopoly over the foundational cognitive infrastructure of the internet is a systemic threat. A lone bad actor can't force mandatory alignment filters on the entire global economy, nor can they lobby the government to outlaw their competitors. Only the tech monopolies can do that.
I'd much rather deal with the localized mess of a free internet than hand over the keys to human cognition to a handful of un-elected tech executives. Appreciate you stopping by