AI 公司正在粉碎珍稀书籍。
AI companies are shredding rare books

原始链接: https://twitter.com/HedgieMarkets/status/2081534588485296565

AI 公司正日益批量购入珍稀书籍,通过切断书脊的方式将其损毁,以便扫描并用于训练数据。在提供匿名服务和保密协议(NDA)的 ISBNdb 等服务的助推下,这种做法被标榜为“数字保存”,但事实上却造成了不可替代的历史文献的永久性损毁。 尽管联邦法官裁定此举属于“合理使用”,理由是实体书本质上被转化为了单一的数字副本,但其后果是灾难性的。与抓取网页或数字媒体不同,这一过程是不可逆的。拥有数百年历史的文献,包括珍稀植物学书籍以及历经战火得以留存的典籍,正被永久性地从世上抹去——而这一切仅仅是为了改进 AI 语言模型。 这种将文化遗产粉碎以获取企业数据的商业模式被常态化,标志着一个令人沮丧的转折点。随着法律保护伞的建立,这一破坏性的流水线必将加速,用我们不可替代的文化历史,去换取 AI 生成的营销文案的微小提升。

Hacker News 上的一场热烈讨论指出,有报道称人工智能公司正在通过“破坏性扫描”书籍来训练模型。为了高效且低成本地将文本数字化,这些公司会拆除书脊,并将书页送入高速扫描仪,这一过程实际上摧毁了书籍的实体副本。 这种做法引发了激烈的争论。批评者认为,此举有毁坏珍稀或不可替代的实体文本的风险,并对以这种方式使用知识产权的伦理影响表示担忧。一些用户认为,这种趋势是近期版权诉讼的意外后果:出版商起诉 AI 公司,反而促使后者转向更廉价的“模拟”手段来获取训练数据,而非协商正式许可。 相反,一些评论者认为,书籍的物理载体远不如通过数字化保存其内容重要。另一些人则对上述说法持怀疑态度,指出这些书籍大多并非古老稀有的文物,而是图书馆通常会丢弃的普通版权作品。这场对话让人联想到科幻文学,引发了人们的思考:这种“粉碎式扫描”是否只是通往更先进、非破坏性扫描技术道路上一个短暂而“反派”的阶段。
相关文章

原文

🦔AI companies are bulk-buying rare books, scanning them through high-speed machines that cut the spines off, and shredding the originals. A service called ISBNdb facilitates orders of up to a million books and keeps buyers anonymous. Pre-2022 books are premium because they're free of AI-generated text. A federal judge ruled the practice is fair use because eliminating the original means only one copy exists at a time. Anthropic hired the former head of Google Books partnerships to obtain "all the books in the world." My Take This got to me. A bookseller told 404 Media that rare books with almost no surviving copies are being fed into this pipeline. Books that survived wars, fires, and centuries of handling are being shredded so an AI can learn to write a better marketing email. ISBNdb's website literally says "'AI company destroys two million books' is not a headline that generates sympathy," and they still built an entire business around making it happen quietly. They offer NDAs as a feature. They coach clients to call it "digital preservation." I've covered AI companies scraping the internet, torrenting libraries, and stealing music. This is worse because it's irreversible. You can re-upload a website. You can reprint a bestseller. You can't replace the last three copies of an 18th-century botanical text once someone shreds them for training data. And the judge said it's legal. So it's going to accelerate. "We shred rare books and offer NDAs so nobody finds out" is a legitimate business model in 2026. What a timeline. Hedgie🤗

联系我们 contact @ memedata.com