home.social

#model-collapse — Public Fediverse posts

Live and recent posts from across the Fediverse tagged #model-collapse, aggregated by home.social.

fetched live
  1. Helpful to see language put to the kettling of aesthetic and cultural variety:

    "diversity is variance, and optimisers are built to minimise variance. A recommender that’s a mediocre predictor of what you’d love in a rich, diverse world becomes an excellent predictor of what you’ll click in the impoverished one it creates."
    laurenleek.substack.com/p/temp

    When people question San Francisco's limit on chain franchises, I wander over to the Toronado or Molotovs or the NocNoc or the Page or the Peacock or Cafe International and there I easily ignore them.

    #ModelCollapse #overfitting

  2. Helpful to see language put to the kettling of aesthetic and cultural variety:

    "diversity is variance, and optimisers are built to minimise variance. A recommender that’s a mediocre predictor of what you’d love in a rich, diverse world becomes an excellent predictor of what you’ll click in the impoverished one it creates."
    laurenleek.substack.com/p/temp

    When people question San Francisco's limit on chain franchises, I wander over to the Toronado or Molotovs or the NocNoc or the Page or the Peacock or Cafe International and there I easily ignore them.

    #ModelCollapse #overfitting

  3. Helpful to see language put to the kettling of aesthetic and cultural variety:

    "diversity is variance, and optimisers are built to minimise variance. A recommender that’s a mediocre predictor of what you’d love in a rich, diverse world becomes an excellent predictor of what you’ll click in the impoverished one it creates."
    laurenleek.substack.com/p/temp

    When people question San Francisco's limit on chain franchises, I wander over to the Toronado or Molotovs or the NocNoc or the Page or the Peacock or Cafe International and there I easily ignore them.

    #ModelCollapse #overfitting

  4. Helpful to see language put to the kettling of aesthetic and cultural variety:

    "diversity is variance, and optimisers are built to minimise variance. A recommender that’s a mediocre predictor of what you’d love in a rich, diverse world becomes an excellent predictor of what you’ll click in the impoverished one it creates."
    laurenleek.substack.com/p/temp

    When people question San Francisco's limit on chain franchises, I wander over to the Toronado or Molotovs or the NocNoc or the Page or the Peacock or Cafe International and there I easily ignore them.

    #ModelCollapse #overfitting

  5. AI zatruła internet. Teraz firmy masowo skupują papierowe książki, by ratować swoje modele

    Sztuczna inteligencja wpadła we własną pułapkę.

    Generatory tekstu zalały sieć marną jakościowo, masową produkcją treści, tworząc cyfrowe wysypisko śmieci. Efekt? Internet przestaje być wystarczająco wiarygodnym źródłem danych treningowych dla kolejnych generacji modeli językowych. Czerpanie wiedzy ze skażonej sieci zwiększa ryzyko tzw. zapaści modelu (model collapse), w której algorytmy uczą się na własnych błędach, stając się wtórne i podatne na halucynacje.

    Aby powstrzymać tę degradację, giganci technologiczni wykonali gwałtowny zwrot ku tradycji. Drukowane książki wydane przed 2022 rokiem stały się nowym złotem branży IT, bo gwarantują dostęp do wiedzy wytworzonej wyłącznie przez ludzi.

    ISBNdb i dyskretny skup na masową skalę

    Kluczowym graczem w tym procesie stał się serwis ISBNdb, dysponujący największą na świecie bazą danych o rynku wydawniczym. Podmiot, który dotąd wspierał biblioteki i księgarzy, stworzył osobną usługę pośrednictwa w hurtowym zakupie fizycznych książek dla laboratoriów AI – realizując zamówienia idące w dziesiątki i setki tysięcy egzemplarzy.

    Baza numerów ISBN pozwala inżynierom precyzyjnie przeczesywać rynek wtórny i unikać wielokrotnego kupowania tych samych tytułów. Automatyczne systemy zakupowe, takie jak AutoBuy w serwisach Alibris czy Biblio, stały się jednym z narzędzi wykorzystywanych przy hurtowym pozyskiwaniu książek, wyłapując całe kategorie wolumenów bez zwracania uwagi na cenę czy ich stan zachowania.

    Cały proces osłonięty jest rygorystycznymi umowami NDA. Firmy technologiczne dbają o zachowanie poufności, ponieważ masowa digitalizacja wiąże się z tzw. skanowaniem destrukcyjnym – gilotynowaniem grzbietów książek, by pojedyncze kartki mogły szybko przejść przez automatyczny skaner.

    Sądowy precedens i ryzyko dla unikalnych wydań

    Przemysłowe niszczenie papieru paradoksalnie zyskało podbudowę prawną. W procesie wytoczonym firmie Anthropic przez pisarzy, sędzia federalny William Alsup orzekł, że cyfryzacja zakupionego tomu mieści się w granicach dozwolonego użytku (fair use) właśnie dlatego, że fizyczny egzemplarz został zniszczony. Cyfrowy plik nie stworzył nowej kopii na rynku, a jedynie zastąpił zniszczony nośnik.

    Anthropic zniszczył miliony książek w celu szkolenia modeli AI

    I choć strona prawna wydaje się zabezpieczona, zjawisko budzi zrozumiały niepokój badaczy i antykwariuszy. Skup nie dotyczy wyłącznie popularnej literatury, ale sięga po rzadkie pozycje naukowe i niszowe wydania akademickie. Masowe niszczenie takich wolumenów w maszynach skanujących może prowadzić do utraty z przestrzeni fizycznej dzieł, które nigdy nie miały i nie będą miały cyfrowych dodruków.

    #Anthropic #dataPoisoning #ISBNdb #literatura #ModelCollapse #rynekKsiążki #sztucznaInteligencja #treningAI
  6. AI zatruła internet. Teraz firmy masowo skupują papierowe książki, by ratować swoje modele

    Sztuczna inteligencja wpadła we własną pułapkę.

    Generatory tekstu zalały sieć marną jakościowo, masową produkcją treści, tworząc cyfrowe wysypisko śmieci. Efekt? Internet przestaje być wystarczająco wiarygodnym źródłem danych treningowych dla kolejnych generacji modeli językowych. Czerpanie wiedzy ze skażonej sieci zwiększa ryzyko tzw. zapaści modelu (model collapse), w której algorytmy uczą się na własnych błędach, stając się wtórne i podatne na halucynacje.

    Aby powstrzymać tę degradację, giganci technologiczni wykonali gwałtowny zwrot ku tradycji. Drukowane książki wydane przed 2022 rokiem stały się nowym złotem branży IT, bo gwarantują dostęp do wiedzy wytworzonej wyłącznie przez ludzi.

    ISBNdb i dyskretny skup na masową skalę

    Kluczowym graczem w tym procesie stał się serwis ISBNdb, dysponujący największą na świecie bazą danych o rynku wydawniczym. Podmiot, który dotąd wspierał biblioteki i księgarzy, stworzył osobną usługę pośrednictwa w hurtowym zakupie fizycznych książek dla laboratoriów AI – realizując zamówienia idące w dziesiątki i setki tysięcy egzemplarzy.

    Baza numerów ISBN pozwala inżynierom precyzyjnie przeczesywać rynek wtórny i unikać wielokrotnego kupowania tych samych tytułów. Automatyczne systemy zakupowe, takie jak AutoBuy w serwisach Alibris czy Biblio, stały się jednym z narzędzi wykorzystywanych przy hurtowym pozyskiwaniu książek, wyłapując całe kategorie wolumenów bez zwracania uwagi na cenę czy ich stan zachowania.

    Cały proces osłonięty jest rygorystycznymi umowami NDA. Firmy technologiczne dbają o zachowanie poufności, ponieważ masowa digitalizacja wiąże się z tzw. skanowaniem destrukcyjnym – gilotynowaniem grzbietów książek, by pojedyncze kartki mogły szybko przejść przez automatyczny skaner.

    Sądowy precedens i ryzyko dla unikalnych wydań

    Przemysłowe niszczenie papieru paradoksalnie zyskało podbudowę prawną. W procesie wytoczonym firmie Anthropic przez pisarzy, sędzia federalny William Alsup orzekł, że cyfryzacja zakupionego tomu mieści się w granicach dozwolonego użytku (fair use) właśnie dlatego, że fizyczny egzemplarz został zniszczony. Cyfrowy plik nie stworzył nowej kopii na rynku, a jedynie zastąpił zniszczony nośnik.

    Anthropic zniszczył miliony książek w celu szkolenia modeli AI

    I choć strona prawna wydaje się zabezpieczona, zjawisko budzi zrozumiały niepokój badaczy i antykwariuszy. Skup nie dotyczy wyłącznie popularnej literatury, ale sięga po rzadkie pozycje naukowe i niszowe wydania akademickie. Masowe niszczenie takich wolumenów w maszynach skanujących może prowadzić do utraty z przestrzeni fizycznej dzieł, które nigdy nie miały i nie będą miały cyfrowych dodruków.

    #Anthropic #dataPoisoning #ISBNdb #literatura #ModelCollapse #rynekKsiążki #sztucznaInteligencja #treningAI
  7. AI zatruła internet. Teraz firmy masowo skupują papierowe książki, by ratować swoje modele

    Sztuczna inteligencja wpadła we własną pułapkę.

    Generatory tekstu zalały sieć marną jakościowo, masową produkcją treści, tworząc cyfrowe wysypisko śmieci. Efekt? Internet przestaje być wystarczająco wiarygodnym źródłem danych treningowych dla kolejnych generacji modeli językowych. Czerpanie wiedzy ze skażonej sieci zwiększa ryzyko tzw. zapaści modelu (model collapse), w której algorytmy uczą się na własnych błędach, stając się wtórne i podatne na halucynacje.

    Aby powstrzymać tę degradację, giganci technologiczni wykonali gwałtowny zwrot ku tradycji. Drukowane książki wydane przed 2022 rokiem stały się nowym złotem branży IT, bo gwarantują dostęp do wiedzy wytworzonej wyłącznie przez ludzi.

    ISBNdb i dyskretny skup na masową skalę

    Kluczowym graczem w tym procesie stał się serwis ISBNdb, dysponujący największą na świecie bazą danych o rynku wydawniczym. Podmiot, który dotąd wspierał biblioteki i księgarzy, stworzył osobną usługę pośrednictwa w hurtowym zakupie fizycznych książek dla laboratoriów AI – realizując zamówienia idące w dziesiątki i setki tysięcy egzemplarzy.

    Baza numerów ISBN pozwala inżynierom precyzyjnie przeczesywać rynek wtórny i unikać wielokrotnego kupowania tych samych tytułów. Automatyczne systemy zakupowe, takie jak AutoBuy w serwisach Alibris czy Biblio, stały się jednym z narzędzi wykorzystywanych przy hurtowym pozyskiwaniu książek, wyłapując całe kategorie wolumenów bez zwracania uwagi na cenę czy ich stan zachowania.

    Cały proces osłonięty jest rygorystycznymi umowami NDA. Firmy technologiczne dbają o zachowanie poufności, ponieważ masowa digitalizacja wiąże się z tzw. skanowaniem destrukcyjnym – gilotynowaniem grzbietów książek, by pojedyncze kartki mogły szybko przejść przez automatyczny skaner.

    Sądowy precedens i ryzyko dla unikalnych wydań

    Przemysłowe niszczenie papieru paradoksalnie zyskało podbudowę prawną. W procesie wytoczonym firmie Anthropic przez pisarzy, sędzia federalny William Alsup orzekł, że cyfryzacja zakupionego tomu mieści się w granicach dozwolonego użytku (fair use) właśnie dlatego, że fizyczny egzemplarz został zniszczony. Cyfrowy plik nie stworzył nowej kopii na rynku, a jedynie zastąpił zniszczony nośnik.

    Anthropic zniszczył miliony książek w celu szkolenia modeli AI

    I choć strona prawna wydaje się zabezpieczona, zjawisko budzi zrozumiały niepokój badaczy i antykwariuszy. Skup nie dotyczy wyłącznie popularnej literatury, ale sięga po rzadkie pozycje naukowe i niszowe wydania akademickie. Masowe niszczenie takich wolumenów w maszynach skanujących może prowadzić do utraty z przestrzeni fizycznej dzieł, które nigdy nie miały i nie będą miały cyfrowych dodruków.

    #Anthropic #dataPoisoning #ISBNdb #literatura #ModelCollapse #rynekKsiążki #sztucznaInteligencja #treningAI
  8. AI zatruła internet. Teraz firmy masowo skupują papierowe książki, by ratować swoje modele

    Sztuczna inteligencja wpadła we własną pułapkę.

    Generatory tekstu zalały sieć marną jakościowo, masową produkcją treści, tworząc cyfrowe wysypisko śmieci. Efekt? Internet przestaje być wystarczająco wiarygodnym źródłem danych treningowych dla kolejnych generacji modeli językowych. Czerpanie wiedzy ze skażonej sieci zwiększa ryzyko tzw. zapaści modelu (model collapse), w której algorytmy uczą się na własnych błędach, stając się wtórne i podatne na halucynacje.

    Aby powstrzymać tę degradację, giganci technologiczni wykonali gwałtowny zwrot ku tradycji. Drukowane książki wydane przed 2022 rokiem stały się nowym złotem branży IT, bo gwarantują dostęp do wiedzy wytworzonej wyłącznie przez ludzi.

    ISBNdb i dyskretny skup na masową skalę

    Kluczowym graczem w tym procesie stał się serwis ISBNdb, dysponujący największą na świecie bazą danych o rynku wydawniczym. Podmiot, który dotąd wspierał biblioteki i księgarzy, stworzył osobną usługę pośrednictwa w hurtowym zakupie fizycznych książek dla laboratoriów AI – realizując zamówienia idące w dziesiątki i setki tysięcy egzemplarzy.

    Baza numerów ISBN pozwala inżynierom precyzyjnie przeczesywać rynek wtórny i unikać wielokrotnego kupowania tych samych tytułów. Automatyczne systemy zakupowe, takie jak AutoBuy w serwisach Alibris czy Biblio, stały się jednym z narzędzi wykorzystywanych przy hurtowym pozyskiwaniu książek, wyłapując całe kategorie wolumenów bez zwracania uwagi na cenę czy ich stan zachowania.

    Cały proces osłonięty jest rygorystycznymi umowami NDA. Firmy technologiczne dbają o zachowanie poufności, ponieważ masowa digitalizacja wiąże się z tzw. skanowaniem destrukcyjnym – gilotynowaniem grzbietów książek, by pojedyncze kartki mogły szybko przejść przez automatyczny skaner.

    Sądowy precedens i ryzyko dla unikalnych wydań

    Przemysłowe niszczenie papieru paradoksalnie zyskało podbudowę prawną. W procesie wytoczonym firmie Anthropic przez pisarzy, sędzia federalny William Alsup orzekł, że cyfryzacja zakupionego tomu mieści się w granicach dozwolonego użytku (fair use) właśnie dlatego, że fizyczny egzemplarz został zniszczony. Cyfrowy plik nie stworzył nowej kopii na rynku, a jedynie zastąpił zniszczony nośnik.

    Anthropic zniszczył miliony książek w celu szkolenia modeli AI

    I choć strona prawna wydaje się zabezpieczona, zjawisko budzi zrozumiały niepokój badaczy i antykwariuszy. Skup nie dotyczy wyłącznie popularnej literatury, ale sięga po rzadkie pozycje naukowe i niszowe wydania akademickie. Masowe niszczenie takich wolumenów w maszynach skanujących może prowadzić do utraty z przestrzeni fizycznej dzieł, które nigdy nie miały i nie będą miały cyfrowych dodruków.

    #Anthropic #dataPoisoning #ISBNdb #literatura #ModelCollapse #rynekKsiążki #sztucznaInteligencja #treningAI
  9. “As AI companies search for more training data to improve their models, one company is offering old, printed books as an ideal source because they are guaranteed to be free of the very AI slop AI companies are producing.

    “The world's best AI training data is sitting on a shelf,” ISBNdb, a company that produces what it claims is “the world’s largest book database,” and that offers high-volume book acquisition services for AI companies, says on its site. “Books represent curated, peer-reviewed, domain-specific human knowledge, structured in a way no web crawl can replicate. Dense, edited, authoritative.”

    In one article on its site, ISBNdb explains that printed books published before 2022 are ideal for AI training data because they don’t include AI generated text. As the article correctly notes, much of the data that AI companies can scrape from the internet today is likely to include AI generated text, which could result in “model collapse,” a process by which AI models that are trained on AI generated data results in worse models that are more prone to errors. The article also notes that book authors who object to their writing being scraped for training purposes can now easily poison AI models by producing writing designed to manipulate and sabotage the resulting AI models.”

    404media.co/ai-companies-are-b

    #AI #GenerativeAI #AISlop #Books #BigTech #ModelCollapse

  10. “As AI companies search for more training data to improve their models, one company is offering old, printed books as an ideal source because they are guaranteed to be free of the very AI slop AI companies are producing.

    “The world's best AI training data is sitting on a shelf,” ISBNdb, a company that produces what it claims is “the world’s largest book database,” and that offers high-volume book acquisition services for AI companies, says on its site. “Books represent curated, peer-reviewed, domain-specific human knowledge, structured in a way no web crawl can replicate. Dense, edited, authoritative.”

    In one article on its site, ISBNdb explains that printed books published before 2022 are ideal for AI training data because they don’t include AI generated text. As the article correctly notes, much of the data that AI companies can scrape from the internet today is likely to include AI generated text, which could result in “model collapse,” a process by which AI models that are trained on AI generated data results in worse models that are more prone to errors. The article also notes that book authors who object to their writing being scraped for training purposes can now easily poison AI models by producing writing designed to manipulate and sabotage the resulting AI models.”

    404media.co/ai-companies-are-b

    #AI #GenerativeAI #AISlop #Books #BigTech #ModelCollapse

  11. “As AI companies search for more training data to improve their models, one company is offering old, printed books as an ideal source because they are guaranteed to be free of the very AI slop AI companies are producing.

    “The world's best AI training data is sitting on a shelf,” ISBNdb, a company that produces what it claims is “the world’s largest book database,” and that offers high-volume book acquisition services for AI companies, says on its site. “Books represent curated, peer-reviewed, domain-specific human knowledge, structured in a way no web crawl can replicate. Dense, edited, authoritative.”

    In one article on its site, ISBNdb explains that printed books published before 2022 are ideal for AI training data because they don’t include AI generated text. As the article correctly notes, much of the data that AI companies can scrape from the internet today is likely to include AI generated text, which could result in “model collapse,” a process by which AI models that are trained on AI generated data results in worse models that are more prone to errors. The article also notes that book authors who object to their writing being scraped for training purposes can now easily poison AI models by producing writing designed to manipulate and sabotage the resulting AI models.”

    404media.co/ai-companies-are-b

    #AI #GenerativeAI #AISlop #Books #BigTech #ModelCollapse

  12. “As AI companies search for more training data to improve their models, one company is offering old, printed books as an ideal source because they are guaranteed to be free of the very AI slop AI companies are producing.

    “The world's best AI training data is sitting on a shelf,” ISBNdb, a company that produces what it claims is “the world’s largest book database,” and that offers high-volume book acquisition services for AI companies, says on its site. “Books represent curated, peer-reviewed, domain-specific human knowledge, structured in a way no web crawl can replicate. Dense, edited, authoritative.”

    In one article on its site, ISBNdb explains that printed books published before 2022 are ideal for AI training data because they don’t include AI generated text. As the article correctly notes, much of the data that AI companies can scrape from the internet today is likely to include AI generated text, which could result in “model collapse,” a process by which AI models that are trained on AI generated data results in worse models that are more prone to errors. The article also notes that book authors who object to their writing being scraped for training purposes can now easily poison AI models by producing writing designed to manipulate and sabotage the resulting AI models.”

    404media.co/ai-companies-are-b

    #AI #GenerativeAI #AISlop #Books #BigTech #ModelCollapse

  13. “As AI companies search for more training data to improve their models, one company is offering old, printed books as an ideal source because they are guaranteed to be free of the very AI slop AI companies are producing.

    “The world's best AI training data is sitting on a shelf,” ISBNdb, a company that produces what it claims is “the world’s largest book database,” and that offers high-volume book acquisition services for AI companies, says on its site. “Books represent curated, peer-reviewed, domain-specific human knowledge, structured in a way no web crawl can replicate. Dense, edited, authoritative.”

    In one article on its site, ISBNdb explains that printed books published before 2022 are ideal for AI training data because they don’t include AI generated text. As the article correctly notes, much of the data that AI companies can scrape from the internet today is likely to include AI generated text, which could result in “model collapse,” a process by which AI models that are trained on AI generated data results in worse models that are more prone to errors. The article also notes that book authors who object to their writing being scraped for training purposes can now easily poison AI models by producing writing designed to manipulate and sabotage the resulting AI models.”

    404media.co/ai-companies-are-b

    #AI #GenerativeAI #AISlop #Books #BigTech #ModelCollapse

  14. Yay! My essay on the impacts of Large Language Model (#LLM) #AI in #archaeology was just published:

    doi.org/10.11141/ia.71.15

    It looks at #bots and mass scraping on the infrastructure supporting #opendata and #openaccess. It also looks at the incentives that encourage the mass-production of #bullshit that may lead to #modelcollapse but more likely less dramatic and more dreary outcomes.

    Enjoy!
    #stochasticparrots #digitalhumanities #openscience

  15. Yay! My essay on the impacts of Large Language Model (#LLM) #AI in #archaeology was just published:

    doi.org/10.11141/ia.71.15

    It looks at #bots and mass scraping on the infrastructure supporting #opendata and #openaccess. It also looks at the incentives that encourage the mass-production of #bullshit that may lead to #modelcollapse but more likely less dramatic and more dreary outcomes.

    Enjoy!
    #stochasticparrots #digitalhumanities #openscience

  16. Yay! My essay on the impacts of Large Language Model (#LLM) #AI in #archaeology was just published:

    doi.org/10.11141/ia.71.15

    It looks at #bots and mass scraping on the infrastructure supporting #opendata and #openaccess. It also looks at the incentives that encourage the mass-production of #bullshit that may lead to #modelcollapse but more likely less dramatic and more dreary outcomes.

    Enjoy!
    #stochasticparrots #digitalhumanities #openscience

  17. Yay! My essay on the impacts of Large Language Model (#LLM) #AI in #archaeology was just published:

    doi.org/10.11141/ia.71.15

    It looks at #bots and mass scraping on the infrastructure supporting #opendata and #openaccess. It also looks at the incentives that encourage the mass-production of #bullshit that may lead to #modelcollapse but more likely less dramatic and more dreary outcomes.

    Enjoy!
    #stochasticparrots #digitalhumanities #openscience

  18. @ApostateEnglishman This spelling appears in some dictionaries because it was used in 1646. As far as I can tell it was used only _once_. This misspelling of "loyalty" was probably a typo or mistranslation by the original author. Or, it may be an error introduced in more recent times during scanning/OCR. I haven't seen a photo of the original page so I can't confirm, but I have seen this sort of glitch happen.

    Nevertheless, it's bizarre to include such a rare and archaic word in spell-check dictionaries!

    How did this happen? I think it may be a consequence of LLMs scraping content from online sources, using what it finds without the intelligence to discern between quality and slop, and negligent humans failing to review machine-generated content before declaring "LGTM, ship it!" Next, that LLM gets scraped by other LLMs, which indiscriminately incorporate the errors into their own AI model training corpus in an ever-worsening "Habsburg AI" feedback loop.

    Thus, it seems one person's typo nearly 400 years ago has resurfaced and is contributing to AI Model Collapse.

    #AI #LLM #LLMs #AISlop #HabsburgAI #AIModelCollapse #ModelCollapse #AutoCarrot

  19. @ApostateEnglishman This spelling appears in some dictionaries because it was used in 1646. As far as I can tell it was used only _once_. This misspelling of "loyalty" was probably a typo or mistranslation by the original author. Or, it may be an error introduced in more recent times during scanning/OCR. I haven't seen a photo of the original page so I can't confirm, but I have seen this sort of glitch happen.

    Nevertheless, it's bizarre to include such a rare and archaic word in spell-check dictionaries!

    How did this happen? I think it may be a consequence of LLMs scraping content from online sources, using what it finds without the intelligence to discern between quality and slop, and negligent humans failing to review machine-generated content before declaring "LGTM, ship it!" Next, that LLM gets scraped by other LLMs, which indiscriminately incorporate the errors into their own AI model training corpus in an ever-worsening "Habsburg AI" feedback loop.

    Thus, it seems one person's typo nearly 400 years ago has resurfaced and is contributing to AI Model Collapse.

    #AI #LLM #LLMs #AISlop #HabsburgAI #AIModelCollapse #ModelCollapse #AutoCarrot

  20. @ApostateEnglishman This spelling appears in some dictionaries because it was used in 1646. As far as I can tell it was used only _once_. This misspelling of "loyalty" was probably a typo or mistranslation by the original author. Or, it may be an error introduced in more recent times during scanning/OCR. I haven't seen a photo of the original page so I can't confirm, but I have seen this sort of glitch happen.

    Nevertheless, it's bizarre to include such a rare and archaic word in spell-check dictionaries!

    How did this happen? I think it may be a consequence of LLMs scraping content from online sources, using what it finds without the intelligence to discern between quality and slop, and negligent humans failing to review machine-generated content before declaring "LGTM, ship it!" Next, that LLM gets scraped by other LLMs, which indiscriminately incorporate the errors into their own AI model training corpus in an ever-worsening "Habsburg AI" feedback loop.

    Thus, it seems one person's typo nearly 400 years ago has resurfaced and is contributing to AI Model Collapse.

    #AI #LLM #LLMs #AISlop #HabsburgAI #AIModelCollapse #ModelCollapse #AutoCarrot

  21. @ApostateEnglishman This spelling appears in some dictionaries because it was used in 1646. As far as I can tell it was used only _once_. This misspelling of "loyalty" was probably a typo or mistranslation by the original author. Or, it may be an error introduced in more recent times during scanning/OCR. I haven't seen a photo of the original page so I can't confirm, but I have seen this sort of glitch happen.

    Nevertheless, it's bizarre to include such a rare and archaic word in spell-check dictionaries!

    How did this happen? I think it may be a consequence of LLMs scraping content from online sources, using what it finds without the intelligence to discern between quality and slop, and negligent humans failing to review machine-generated content before declaring "LGTM, ship it!" Next, that LLM gets scraped by other LLMs, which indiscriminately incorporate the errors into their own AI model training corpus in an ever-worsening "Habsburg AI" feedback loop.

    Thus, it seems one person's typo nearly 400 years ago has resurfaced and is contributing to AI Model Collapse.

    #AI #LLM #LLMs #AISlop #HabsburgAI #AIModelCollapse #ModelCollapse #AutoCarrot

  22. @ApostateEnglishman This spelling appears in some dictionaries because it was used in 1646. As far as I can tell it was used only _once_. This misspelling of "loyalty" was probably a typo or mistranslation by the original author. Or, it may be an error introduced in more recent times during scanning/OCR. I haven't seen a photo of the original page so I can't confirm, but I have seen this sort of glitch happen.

    Nevertheless, it's bizarre to include such a rare and archaic word in spell-check dictionaries!

    How did this happen? I think it may be a consequence of LLMs scraping content from online sources, using what it finds without the intelligence to discern between quality and slop, and negligent humans failing to review machine-generated content before declaring "LGTM, ship it!" Next, that LLM gets scraped by other LLMs, which indiscriminately incorporate the errors into their own AI model training corpus in an ever-worsening "Habsburg AI" feedback loop.

    Thus, it seems one person's typo nearly 400 years ago has resurfaced and is contributing to AI Model Collapse.

    #AI #LLM #LLMs #AISlop #HabsburgAI #AIModelCollapse #ModelCollapse #AutoCarrot

  23. À force d’utiliser l’#IA, les #journalistes risquent-ils d’appauvrir la langue ?
    theconversation.com/a-force-du
    Quand les systèmes commencent à être entraînés à partir de textes produits par d’autres IA arrive le #modelcollapse ou #effondrement du modèle un processus de #dégénérescence où les données générées par un modèle finissent par contaminer l’entraînement des générations suivantes.
    + Il y a de textes artificiels - les modèles sont exposés à la diversité réelle des usages humains de la langue

  24. À force d’utiliser l’#IA, les #journalistes risquent-ils d’appauvrir la langue ?
    theconversation.com/a-force-du
    Quand les systèmes commencent à être entraînés à partir de textes produits par d’autres IA arrive le #modelcollapse ou #effondrement du modèle un processus de #dégénérescence où les données générées par un modèle finissent par contaminer l’entraînement des générations suivantes.
    + Il y a de textes artificiels - les modèles sont exposés à la diversité réelle des usages humains de la langue

  25. À force d’utiliser l’#IA, les #journalistes risquent-ils d’appauvrir la langue ?
    theconversation.com/a-force-du
    Quand les systèmes commencent à être entraînés à partir de textes produits par d’autres IA arrive le #modelcollapse ou #effondrement du modèle un processus de #dégénérescence où les données générées par un modèle finissent par contaminer l’entraînement des générations suivantes.
    + Il y a de textes artificiels - les modèles sont exposés à la diversité réelle des usages humains de la langue

  26. À force d’utiliser l’#IA, les #journalistes risquent-ils d’appauvrir la langue ?
    theconversation.com/a-force-du
    Quand les systèmes commencent à être entraînés à partir de textes produits par d’autres IA arrive le #modelcollapse ou #effondrement du modèle un processus de #dégénérescence où les données générées par un modèle finissent par contaminer l’entraînement des générations suivantes.
    + Il y a de textes artificiels - les modèles sont exposés à la diversité réelle des usages humains de la langue

  27. À force d’utiliser l’#IA, les #journalistes risquent-ils d’appauvrir la langue ?
    theconversation.com/a-force-du
    Quand les systèmes commencent à être entraînés à partir de textes produits par d’autres IA arrive le #modelcollapse ou #effondrement du modèle un processus de #dégénérescence où les données générées par un modèle finissent par contaminer l’entraînement des générations suivantes.
    + Il y a de textes artificiels - les modèles sont exposés à la diversité réelle des usages humains de la langue

  28. 🔴 LIVE NOW ON VORTEX
    📻 Vortex Night ⛓️ (Industrial metal)
    ──────────────
    🎵 MODEL COLLAPSE - SILENT PATH

    ▶️ Écouter / Listen : VorteX [Radio]
    lesonduvortex.net

    💬 Join us on Discord:
    discord.gg/d82hJZBeDE

    #VortexWave #ModelCollapse #Ambient #Post-Rock #2000s

  29. 🔴 LIVE NOW ON VORTEX
    📻 Vortex Night ⛓️ (Industrial metal)
    ──────────────
    🎵 MODEL COLLAPSE - SILENT PATH

    ▶️ Écouter / Listen : VorteX [Radio]
    lesonduvortex.net

    💬 Join us on Discord:
    discord.gg/d82hJZBeDE

    #VortexWave #ModelCollapse #Ambient #Post-Rock #2000s

  30. 🔴 LIVE NOW ON VORTEX
    📻 Vortex Night ⛓️ (Industrial metal)
    ──────────────
    🎵 MODEL COLLAPSE - SILENT PATH

    ▶️ Écouter / Listen : VorteX [Radio]
    lesonduvortex.net

    💬 Join us on Discord:
    discord.gg/d82hJZBeDE

    #VortexWave #ModelCollapse #Ambient #Post-Rock #2000s

  31. 🔴 LIVE NOW ON VORTEX
    📻 Vortex Night ⛓️ (Industrial metal)
    ──────────────
    🎵 MODEL COLLAPSE - SILENT PATH

    ▶️ Écouter / Listen : VorteX [Radio]
    lesonduvortex.net

    💬 Join us on Discord:
    discord.gg/d82hJZBeDE

    #VortexWave #ModelCollapse #Ambient #Post-Rock #2000s

  32. 🔴 LIVE NOW ON VORTEX
    📻 Vortex Night ⛓️ (Industrial metal)
    ──────────────
    🎵 MODEL COLLAPSE - SILENT PATH

    ▶️ Écouter / Listen : VorteX [Radio]
    lesonduvortex.net

    💬 Join us on Discord:
    discord.gg/d82hJZBeDE

    #VortexWave #ModelCollapse #Ambient #Post-Rock #2000s

  33. AI makes mistakes – I still notice them because I have prior knowledge.

    But what about young people who use AI as their primary source of information?

    And: what happens when this generation trains the next AI – with the knowledge they got from AI?

    Does ignorance compound itself?

    #ai #modelcollapse #ailiteracy #education

  34. 🟩 𝗘𝗫𝗛𝗜𝗕𝗜𝗧𝗜𝗢𝗡: 𝐿𝑎𝑡𝑒𝑛𝑡 𝑆𝑝𝑎𝑐𝑒
    1–30 April | Aksioma Project Space
    ❕ 𝗢𝗽𝗲𝗻𝗶𝗻𝗴: 1 April at 8 PM

    In her installation, artist #FelicityHammond offers a speculative glimpse into a not-too-distant future where this new approach to space-based computation has become the dominant position in the AI industry. However, the system continues to battle with the effects of #modelcollapse...

    > aksioma.org/becomingimage/exhi

  35. 🟩 𝗘𝗫𝗛𝗜𝗕𝗜𝗧𝗜𝗢𝗡: 𝐿𝑎𝑡𝑒𝑛𝑡 𝑆𝑝𝑎𝑐𝑒
    1–30 April | Aksioma Project Space
    ❕ 𝗢𝗽𝗲𝗻𝗶𝗻𝗴: 1 April at 8 PM

    In her installation, artist #FelicityHammond offers a speculative glimpse into a not-too-distant future where this new approach to space-based computation has become the dominant position in the AI industry. However, the system continues to battle with the effects of #modelcollapse...

    > aksioma.org/becomingimage/exhi