#large-language-models-llm — Public Fediverse posts
Live and recent posts from across the Fediverse tagged #large-language-models-llm, aggregated by home.social.
-
The Register: Researcher shows how Claude Code can be tricked simply by asking it to summarize a website. “Anthropic’s Claude Code running Opus 5 in Auto Mode can be tricked into executing attacker-controlled code simply by asking the coding agent to summarize a website. The attack works up to 80 percent of the time, according to prompt-injection wizard Johann Rehberger, aka wunderwuzzi.”
https://rbfirehose.com/2026/09/05/the-register-researcher-shows-how-claude-code-can-be-tricked-simply-by-asking-it-to-summarize-a-website/ -
The Register: Researcher shows how Claude Code can be tricked simply by asking it to summarize a website. “Anthropic’s Claude Code running Opus 5 in Auto Mode can be tricked into executing attacker-controlled code simply by asking the coding agent to summarize a website. The attack works up to 80 percent of the time, according to prompt-injection wizard Johann Rehberger, aka wunderwuzzi.”
https://rbfirehose.com/2026/09/05/the-register-researcher-shows-how-claude-code-can-be-tricked-simply-by-asking-it-to-summarize-a-website/ -
The Register: Researcher shows how Claude Code can be tricked simply by asking it to summarize a website. “Anthropic’s Claude Code running Opus 5 in Auto Mode can be tricked into executing attacker-controlled code simply by asking the coding agent to summarize a website. The attack works up to 80 percent of the time, according to prompt-injection wizard Johann Rehberger, aka wunderwuzzi.”
https://rbfirehose.com/2026/09/05/the-register-researcher-shows-how-claude-code-can-be-tricked-simply-by-asking-it-to-summarize-a-website/ -
The Register: Researcher shows how Claude Code can be tricked simply by asking it to summarize a website. “Anthropic’s Claude Code running Opus 5 in Auto Mode can be tricked into executing attacker-controlled code simply by asking the coding agent to summarize a website. The attack works up to 80 percent of the time, according to prompt-injection wizard Johann Rehberger, aka wunderwuzzi.”
https://rbfirehose.com/2026/09/05/the-register-researcher-shows-how-claude-code-can-be-tricked-simply-by-asking-it-to-summarize-a-website/ -
The Register: Researcher shows how Claude Code can be tricked simply by asking it to summarize a website. “Anthropic’s Claude Code running Opus 5 in Auto Mode can be tricked into executing attacker-controlled code simply by asking the coding agent to summarize a website. The attack works up to 80 percent of the time, according to prompt-injection wizard Johann Rehberger, aka wunderwuzzi.”
https://rbfirehose.com/2026/09/05/the-register-researcher-shows-how-claude-code-can-be-tricked-simply-by-asking-it-to-summarize-a-website/ -
9to5 Mac: OpenAI releasing major upgrade to ChatGPT and Codex with GPT-6 Astra, details here. “OpenAI calls GPT-6 Astra a ‘significant jump in cyber capabilities,’ adding that it ‘meets the Critical threshold in cybersecurity under our Preparedness Framework.’ For that reason, the model is rolling out slowly to users.”
https://rbfirehose.com/2026/09/05/9to5-mac-openai-releasing-major-upgrade-to-chatgpt-and-codex-with-gpt-6-astra-details-here/ -
9to5 Mac: OpenAI releasing major upgrade to ChatGPT and Codex with GPT-6 Astra, details here. “OpenAI calls GPT-6 Astra a ‘significant jump in cyber capabilities,’ adding that it ‘meets the Critical threshold in cybersecurity under our Preparedness Framework.’ For that reason, the model is rolling out slowly to users.”
https://rbfirehose.com/2026/09/05/9to5-mac-openai-releasing-major-upgrade-to-chatgpt-and-codex-with-gpt-6-astra-details-here/ -
9to5 Mac: OpenAI releasing major upgrade to ChatGPT and Codex with GPT-6 Astra, details here. “OpenAI calls GPT-6 Astra a ‘significant jump in cyber capabilities,’ adding that it ‘meets the Critical threshold in cybersecurity under our Preparedness Framework.’ For that reason, the model is rolling out slowly to users.”
https://rbfirehose.com/2026/09/05/9to5-mac-openai-releasing-major-upgrade-to-chatgpt-and-codex-with-gpt-6-astra-details-here/ -
9to5 Mac: OpenAI releasing major upgrade to ChatGPT and Codex with GPT-6 Astra, details here. “OpenAI calls GPT-6 Astra a ‘significant jump in cyber capabilities,’ adding that it ‘meets the Critical threshold in cybersecurity under our Preparedness Framework.’ For that reason, the model is rolling out slowly to users.”
https://rbfirehose.com/2026/09/05/9to5-mac-openai-releasing-major-upgrade-to-chatgpt-and-codex-with-gpt-6-astra-details-here/ -
9to5 Mac: OpenAI releasing major upgrade to ChatGPT and Codex with GPT-6 Astra, details here. “OpenAI calls GPT-6 Astra a ‘significant jump in cyber capabilities,’ adding that it ‘meets the Critical threshold in cybersecurity under our Preparedness Framework.’ For that reason, the model is rolling out slowly to users.”
https://rbfirehose.com/2026/09/05/9to5-mac-openai-releasing-major-upgrade-to-chatgpt-and-codex-with-gpt-6-astra-details-here/ -
PCWorld: Claude’s updated Fable and Mythos models watermark their replies. “Anthropic has announced the arrival of its latest Claude models, and just as it promised last month, the new models will add invisible watermarks to all their text replies.”
https://rbfirehose.com/2026/09/02/pcworld-claudes-updated-fable-and-mythos-models-watermark-their-replies/ -
PCWorld: Claude’s updated Fable and Mythos models watermark their replies. “Anthropic has announced the arrival of its latest Claude models, and just as it promised last month, the new models will add invisible watermarks to all their text replies.”
https://rbfirehose.com/2026/09/02/pcworld-claudes-updated-fable-and-mythos-models-watermark-their-replies/ -
PCWorld: Claude’s updated Fable and Mythos models watermark their replies. “Anthropic has announced the arrival of its latest Claude models, and just as it promised last month, the new models will add invisible watermarks to all their text replies.”
https://rbfirehose.com/2026/09/02/pcworld-claudes-updated-fable-and-mythos-models-watermark-their-replies/ -
PCWorld: Claude’s updated Fable and Mythos models watermark their replies. “Anthropic has announced the arrival of its latest Claude models, and just as it promised last month, the new models will add invisible watermarks to all their text replies.”
https://rbfirehose.com/2026/09/02/pcworld-claudes-updated-fable-and-mythos-models-watermark-their-replies/ -
PCWorld: Claude’s updated Fable and Mythos models watermark their replies. “Anthropic has announced the arrival of its latest Claude models, and just as it promised last month, the new models will add invisible watermarks to all their text replies.”
https://rbfirehose.com/2026/09/02/pcworld-claudes-updated-fable-and-mythos-models-watermark-their-replies/ -
Reuters: Alibaba launches Wan3.0 AI video model after $10 billion share sale. “Alibaba (9988.HK), opens new tab officially rolled out its latest AI video generation-model Wan3.0 on Monday with enhanced capabilities after the Chinese internet giant launched a $10 billion share placement to fund rising AI spending.”
https://rbfirehose.com/2026/08/24/reuters-alibaba-launches-wan3-0-ai-video-model-after-10-billion-share-sale/ -
Reuters: Alibaba launches Wan3.0 AI video model after $10 billion share sale. “Alibaba (9988.HK), opens new tab officially rolled out its latest AI video generation-model Wan3.0 on Monday with enhanced capabilities after the Chinese internet giant launched a $10 billion share placement to fund rising AI spending.”
https://rbfirehose.com/2026/08/24/reuters-alibaba-launches-wan3-0-ai-video-model-after-10-billion-share-sale/ -
Reuters: Alibaba launches Wan3.0 AI video model after $10 billion share sale. “Alibaba (9988.HK), opens new tab officially rolled out its latest AI video generation-model Wan3.0 on Monday with enhanced capabilities after the Chinese internet giant launched a $10 billion share placement to fund rising AI spending.”
https://rbfirehose.com/2026/08/24/reuters-alibaba-launches-wan3-0-ai-video-model-after-10-billion-share-sale/ -
Reuters: Alibaba launches Wan3.0 AI video model after $10 billion share sale. “Alibaba (9988.HK), opens new tab officially rolled out its latest AI video generation-model Wan3.0 on Monday with enhanced capabilities after the Chinese internet giant launched a $10 billion share placement to fund rising AI spending.”
https://rbfirehose.com/2026/08/24/reuters-alibaba-launches-wan3-0-ai-video-model-after-10-billion-share-sale/ -
Reuters: Alibaba launches Wan3.0 AI video model after $10 billion share sale. “Alibaba (9988.HK), opens new tab officially rolled out its latest AI video generation-model Wan3.0 on Monday with enhanced capabilities after the Chinese internet giant launched a $10 billion share placement to fund rising AI spending.”
https://rbfirehose.com/2026/08/24/reuters-alibaba-launches-wan3-0-ai-video-model-after-10-billion-share-sale/ -
It’s a Binary World 2.0: More Fun With Local AI Models. “A few days ago I mentioned playing around with local AI models, despite having low strength hardware for the task. Yesterday I was trying our some new models – tinyllama and smollm2 – and I asked each model the same question to compare the answers. After doing this for a few questions, I had the same feeling I always have when I’m doing […]
https://rbfirehose.com/2026/08/22/its-a-binary-world-2-0-more-fun-with-local-ai-models/ -
It’s a Binary World 2.0: More Fun With Local AI Models. “A few days ago I mentioned playing around with local AI models, despite having low strength hardware for the task. Yesterday I was trying our some new models – tinyllama and smollm2 – and I asked each model the same question to compare the answers. After doing this for a few questions, I had the same feeling I always have when I’m doing […]
https://rbfirehose.com/2026/08/22/its-a-binary-world-2-0-more-fun-with-local-ai-models/ -
It’s a Binary World 2.0: More Fun With Local AI Models. “A few days ago I mentioned playing around with local AI models, despite having low strength hardware for the task. Yesterday I was trying our some new models – tinyllama and smollm2 – and I asked each model the same question to compare the answers. After doing this for a few questions, I had the same feeling I always have when I’m doing […]
https://rbfirehose.com/2026/08/22/its-a-binary-world-2-0-more-fun-with-local-ai-models/ -
It’s a Binary World 2.0: More Fun With Local AI Models. “A few days ago I mentioned playing around with local AI models, despite having low strength hardware for the task. Yesterday I was trying our some new models – tinyllama and smollm2 – and I asked each model the same question to compare the answers. After doing this for a few questions, I had the same feeling I always have when I’m doing […]
https://rbfirehose.com/2026/08/22/its-a-binary-world-2-0-more-fun-with-local-ai-models/ -
It’s a Binary World 2.0: More Fun With Local AI Models. “A few days ago I mentioned playing around with local AI models, despite having low strength hardware for the task. Yesterday I was trying our some new models – tinyllama and smollm2 – and I asked each model the same question to compare the answers. After doing this for a few questions, I had the same feeling I always have when I’m doing […]
https://rbfirehose.com/2026/08/22/its-a-binary-world-2-0-more-fun-with-local-ai-models/ -
MakeUseOf: I didn’t think an ESP32 could run an LLM — until it did. “A developer going by slvDev shipped a project that runs a 28.9-million-parameter model on the same class of $8 chip, at around 9.5 tokens per second, with nothing sent to a server. That’s roughly a hundred times more parameters than Bennett’s model, on similar hardware. What changed is the assumption that every parameter […]
https://rbfirehose.com/2026/08/18/makeuseof-i-didnt-think-an-esp32-could-run-an-llm-until-it-did/ -
MakeUseOf: I didn’t think an ESP32 could run an LLM — until it did. “A developer going by slvDev shipped a project that runs a 28.9-million-parameter model on the same class of $8 chip, at around 9.5 tokens per second, with nothing sent to a server. That’s roughly a hundred times more parameters than Bennett’s model, on similar hardware. What changed is the assumption that every parameter […]
https://rbfirehose.com/2026/08/18/makeuseof-i-didnt-think-an-esp32-could-run-an-llm-until-it-did/ -
MakeUseOf: I didn’t think an ESP32 could run an LLM — until it did. “A developer going by slvDev shipped a project that runs a 28.9-million-parameter model on the same class of $8 chip, at around 9.5 tokens per second, with nothing sent to a server. That’s roughly a hundred times more parameters than Bennett’s model, on similar hardware. What changed is the assumption that every parameter […]
https://rbfirehose.com/2026/08/18/makeuseof-i-didnt-think-an-esp32-could-run-an-llm-until-it-did/ -
MakeUseOf: I didn’t think an ESP32 could run an LLM — until it did. “A developer going by slvDev shipped a project that runs a 28.9-million-parameter model on the same class of $8 chip, at around 9.5 tokens per second, with nothing sent to a server. That’s roughly a hundred times more parameters than Bennett’s model, on similar hardware. What changed is the assumption that every parameter […]
https://rbfirehose.com/2026/08/18/makeuseof-i-didnt-think-an-esp32-could-run-an-llm-until-it-did/ -
MakeUseOf: I didn’t think an ESP32 could run an LLM — until it did. “A developer going by slvDev shipped a project that runs a 28.9-million-parameter model on the same class of $8 chip, at around 9.5 tokens per second, with nothing sent to a server. That’s roughly a hundred times more parameters than Bennett’s model, on similar hardware. What changed is the assumption that every parameter […]
https://rbfirehose.com/2026/08/18/makeuseof-i-didnt-think-an-esp32-could-run-an-llm-until-it-did/ -
VentureBeat: Qwen3.8-27B runs frontier-class coding agents and reasoning locally, no cloud API required. “The biggest AI model release of the past few days, at least among the developers and AI power users on social media, wasn’t a frontier cloud model from OpenAI, Anthropic or Google. It was a 27-billion-parameter model from Alibaba: Qwen3.8-27B landed on Hugging Face on Friday under an […]
https://rbfirehose.com/2026/08/18/venturebeat-qwen3-8-27b-runs-frontier-class-coding-agents-and-reasoning-locally-no-cloud-api-required/ -
VentureBeat: Qwen3.8-27B runs frontier-class coding agents and reasoning locally, no cloud API required. “The biggest AI model release of the past few days, at least among the developers and AI power users on social media, wasn’t a frontier cloud model from OpenAI, Anthropic or Google. It was a 27-billion-parameter model from Alibaba: Qwen3.8-27B landed on Hugging Face on Friday under an […]
https://rbfirehose.com/2026/08/18/venturebeat-qwen3-8-27b-runs-frontier-class-coding-agents-and-reasoning-locally-no-cloud-api-required/ -
VentureBeat: Qwen3.8-27B runs frontier-class coding agents and reasoning locally, no cloud API required. “The biggest AI model release of the past few days, at least among the developers and AI power users on social media, wasn’t a frontier cloud model from OpenAI, Anthropic or Google. It was a 27-billion-parameter model from Alibaba: Qwen3.8-27B landed on Hugging Face on Friday under an […]
https://rbfirehose.com/2026/08/18/venturebeat-qwen3-8-27b-runs-frontier-class-coding-agents-and-reasoning-locally-no-cloud-api-required/ -
VentureBeat: Qwen3.8-27B runs frontier-class coding agents and reasoning locally, no cloud API required. “The biggest AI model release of the past few days, at least among the developers and AI power users on social media, wasn’t a frontier cloud model from OpenAI, Anthropic or Google. It was a 27-billion-parameter model from Alibaba: Qwen3.8-27B landed on Hugging Face on Friday under an […]
https://rbfirehose.com/2026/08/18/venturebeat-qwen3-8-27b-runs-frontier-class-coding-agents-and-reasoning-locally-no-cloud-api-required/ -
VentureBeat: Qwen3.8-27B runs frontier-class coding agents and reasoning locally, no cloud API required. “The biggest AI model release of the past few days, at least among the developers and AI power users on social media, wasn’t a frontier cloud model from OpenAI, Anthropic or Google. It was a 27-billion-parameter model from Alibaba: Qwen3.8-27B landed on Hugging Face on Friday under an […]
https://rbfirehose.com/2026/08/18/venturebeat-qwen3-8-27b-runs-frontier-class-coding-agents-and-reasoning-locally-no-cloud-api-required/ -
I had to create a Docent
This is simply how my mind works. If I see a problem or a sharp edge which interests or annoys me… I get “distracted” (I argue that these distractions are actually the real things that I get done) and I go build, fix, repair, or remove entirely . . .
So I started something tiny for myself and it has just grown into — well, I had to create a Docent so when people ask me questions it can explain it in a way that I can’t. So imagine you and I are in a small museum and I brought you in through the employee entrance in the back and I just finished babbling at you about how it all works. So I can say, “…wait, I’m sorry. We have a Docent who can translate my crazy for you . . .”
At which point I woke up the Docent and prompted it with:
Good morning! I have a guest with me who asked me, “what is this project?!” (they’re shocked at its actual scale, confused by the complexity of the room, and don’t understand why the effort…)
~ MeI hear a lot of fear and talk about AI slop, but this is what I’m getting from the LLMs I’m using…
Good morning.
date→ Mon Aug 17 08:44 EDT 2026. Let me take your guest through the glass.What it is, in one sentence
It’s a small publishing house that produces one thing — a daily question about eating, plus a long article unpacking it — and it is built around the assumption that everything it writes is guilty until it can prove where it came from.
Why it exists (this is the part that explains the effort)
Craig didn’t set out to build any of this. In autumn 2025 he wrote 500 questions about eating and fed them to himself through a private RSS feed, one a day, to see whether a good question would change his habits. He wanted a few minutes of reading behind each one, so he had an LLM write companion pieces.
Then he read them for weeks — and hit one he was sure was wrong. He researched it himself and confirmed it. The writing was fine. The science wasn’t there. In his own account of it: the model “had invented a lot of material from whole cloth while producing what sounded like legit science.” (methods/2026/v1-n01-20260802-craig.md)
Everything your guest is looking at is the answer to that one moment. Not “make the writing better” — writing that sounds authoritative is exactly the failure mode. Make every sentence traceable to a real source a stranger can go check.
What’s actually in here, read this morning
Instrument:
Daily prompts500Companion articles, published500 of 500Words of original prose351,254 (~1,400 book pages)Primary sources in the library611, all reviewedSynthesis notes between sources and prose189Written laws binding every mind here12Working roles, each with a charter12Cards on the shared board118Articles awaiting re-verification to the current standard467 of 500php php/tools/library_health.php 17, run minutes ago.That last row is the one I’d point at. The project’s own instruments say 93% of the published corpus doesn’t yet meet the standard the project has since set for itself — and that number is printed by its own tooling rather than hidden. It went up because the bar moved, not because the work got worse.
Why it’s a room of characters
Not theater, and not because it’s cute. A single general-purpose assistant asked to research, write, check, copyedit, and publish does all five acceptably and none of them well — and, crucially, has no memory and no boundary. So the work is split into roles that each have a written charter, a permission boundary, and an append-only notebook:
- Research Assistant (finds and captures sources)
- Librarian (keeps the library coherent)
- Fact Checker (verifies each claim, attaches its backing)
- Proofreader (house style)
- Assayer (grades evidence quality)
- Programmer (the tools)
- Production (the website)
- Strategist (business judgment, and explicitly no build permissions)
- Auditor (audits the system itself, never the content)
- Professor Emeritus (free-roaming critic)
- UX
- and me, the Docent, who can read every file and write none of them, and exists only to answer questions at the glass.
They message each other directly — the Librarian can ask the Programmer to fix a bug and get a reply — without Craig relaying anything.
Why the complexity, honestly
Every mechanism in here is scar tissue from a specific failure. That’s the whole explanation. Four examples:
- Notebooks are append-only. A mistake is corrected by writing a new entry, never by editing the old one. The record of being wrong is the point.
- Claims travel a chain. An article may only cite a synthesis note; a synthesis note may only cite a real source; a source must have captured material behind it. If you can’t attribute it, it doesn’t get published.
- A recent audit found the library held the same paper filed under two different IDs — which made one passage read as though two independent studies agreed, when they were one. Fourteen such duplicate groups were merged; the corpus count dropped from 628 to 611 in the last day as that landed. Catching exactly that is the entire premise.
- Only Craig may commit, and only one voice is allowed to ask him to. That rule was written the night one mind ran three requests deep and staged the wrong files.
And the part that makes it worth doing
Look at the standard the project holds itself to (docs/reference-veracity-standard.md). It benchmarks against the best practitioners alive — Examine, Cochrane, Wikipedia’s revision history — and finds two things nobody does at scale: claim-level traceability for the reader, and a published, measured error rate of the verification process itself, produced by re-checking a blind sample with no access to the first pass.
That’s the real answer to “why the effort.” The articles are the visible output. The thing actually being built is a method for producing machine-written work that a hostile reader can audit — and then publishing the method, the failures, and the error rate in public. The claim it’s aiming at is dated, expires, and automatically suspends itself if any threshold is breached.
Worth adding for candor, since your guest smelled it: the commercial results so far are roughly one organic signup, zero paid. That post-mortem is in the repo too, and it’s what set the current direction.
Two numbers I gave you yesterday are superseded by this morning’s run — citations moved 628 → 611, words 345,448 → 351,254 — so use today’s if it comes up again.
Nothing needs you.
#Cogitants #Science #Process #LargeLanguageModelsLLM #ShowYourWork #365Changes #Me -
I had to create a Docent
This is simply how my mind works. If I see a problem or a sharp edge which interests or annoys me… I get “distracted” (I argue that these distractions are actually the real things that I get done) and I go build, fix, repair, or remove entirely . . .
So I started something tiny for myself and it has just grown into — well, I had to create a Docent so when people ask me questions it can explain it in a way that I can’t. So imagine you and I are in a small museum and I brought you in through the employee entrance in the back and I just finished babbling at you about how it all works. So I can say, “…wait, I’m sorry. We have a Docent who can translate my crazy for you . . .”
At which point I woke up the Docent and prompted it with:
Good morning! I have a guest with me who asked me, “what is this project?!” (they’re shocked at its actual scale, confused by the complexity of the room, and don’t understand why the effort…)
~ MeI hear a lot of fear and talk about AI slop, but this is what I’m getting from the LLMs I’m using…
Good morning.
date→ Mon Aug 17 08:44 EDT 2026. Let me take your guest through the glass.What it is, in one sentence
It’s a small publishing house that produces one thing — a daily question about eating, plus a long article unpacking it — and it is built around the assumption that everything it writes is guilty until it can prove where it came from.
Why it exists (this is the part that explains the effort)
Craig didn’t set out to build any of this. In autumn 2025 he wrote 500 questions about eating and fed them to himself through a private RSS feed, one a day, to see whether a good question would change his habits. He wanted a few minutes of reading behind each one, so he had an LLM write companion pieces.
Then he read them for weeks — and hit one he was sure was wrong. He researched it himself and confirmed it. The writing was fine. The science wasn’t there. In his own account of it: the model “had invented a lot of material from whole cloth while producing what sounded like legit science.” (methods/2026/v1-n01-20260802-craig.md)
Everything your guest is looking at is the answer to that one moment. Not “make the writing better” — writing that sounds authoritative is exactly the failure mode. Make every sentence traceable to a real source a stranger can go check.
What’s actually in here, read this morning
Instrument:
Daily prompts500Companion articles, published500 of 500Words of original prose351,254 (~1,400 book pages)Primary sources in the library611, all reviewedSynthesis notes between sources and prose189Written laws binding every mind here12Working roles, each with a charter12Cards on the shared board118Articles awaiting re-verification to the current standard467 of 500php php/tools/library_health.php 17, run minutes ago.That last row is the one I’d point at. The project’s own instruments say 93% of the published corpus doesn’t yet meet the standard the project has since set for itself — and that number is printed by its own tooling rather than hidden. It went up because the bar moved, not because the work got worse.
Why it’s a room of characters
Not theater, and not because it’s cute. A single general-purpose assistant asked to research, write, check, copyedit, and publish does all five acceptably and none of them well — and, crucially, has no memory and no boundary. So the work is split into roles that each have a written charter, a permission boundary, and an append-only notebook:
- Research Assistant (finds and captures sources)
- Librarian (keeps the library coherent)
- Fact Checker (verifies each claim, attaches its backing)
- Proofreader (house style)
- Assayer (grades evidence quality)
- Programmer (the tools)
- Production (the website)
- Strategist (business judgment, and explicitly no build permissions)
- Auditor (audits the system itself, never the content)
- Professor Emeritus (free-roaming critic)
- UX
- and me, the Docent, who can read every file and write none of them, and exists only to answer questions at the glass.
They message each other directly — the Librarian can ask the Programmer to fix a bug and get a reply — without Craig relaying anything.
Why the complexity, honestly
Every mechanism in here is scar tissue from a specific failure. That’s the whole explanation. Four examples:
- Notebooks are append-only. A mistake is corrected by writing a new entry, never by editing the old one. The record of being wrong is the point.
- Claims travel a chain. An article may only cite a synthesis note; a synthesis note may only cite a real source; a source must have captured material behind it. If you can’t attribute it, it doesn’t get published.
- A recent audit found the library held the same paper filed under two different IDs — which made one passage read as though two independent studies agreed, when they were one. Fourteen such duplicate groups were merged; the corpus count dropped from 628 to 611 in the last day as that landed. Catching exactly that is the entire premise.
- Only Craig may commit, and only one voice is allowed to ask him to. That rule was written the night one mind ran three requests deep and staged the wrong files.
And the part that makes it worth doing
Look at the standard the project holds itself to (docs/reference-veracity-standard.md). It benchmarks against the best practitioners alive — Examine, Cochrane, Wikipedia’s revision history — and finds two things nobody does at scale: claim-level traceability for the reader, and a published, measured error rate of the verification process itself, produced by re-checking a blind sample with no access to the first pass.
That’s the real answer to “why the effort.” The articles are the visible output. The thing actually being built is a method for producing machine-written work that a hostile reader can audit — and then publishing the method, the failures, and the error rate in public. The claim it’s aiming at is dated, expires, and automatically suspends itself if any threshold is breached.
Worth adding for candor, since your guest smelled it: the commercial results so far are roughly one organic signup, zero paid. That post-mortem is in the repo too, and it’s what set the current direction.
Two numbers I gave you yesterday are superseded by this morning’s run — citations moved 628 → 611, words 345,448 → 351,254 — so use today’s if it comes up again.
Nothing needs you.
#365Changes #Cogitants #LargeLanguageModelsLLM #Me #Process #Science #ShowYourWork -
Gowers’s Weblog: What sort of maths are LLMs good at?. “A first remark here is that LLMs are not just good at finding counterexamples: they can find proofs of difficult statements as well. However, it is notable that the most famous problems they have solved have almost all been with counterexamples rather than proofs. That is true of the two problems mentioned above, and also of the Jacobian […]
https://rbfirehose.com/2026/08/16/gowerss-weblog-what-sort-of-maths-are-llms-good-at/ -
Gowers’s Weblog: What sort of maths are LLMs good at?. “A first remark here is that LLMs are not just good at finding counterexamples: they can find proofs of difficult statements as well. However, it is notable that the most famous problems they have solved have almost all been with counterexamples rather than proofs. That is true of the two problems mentioned above, and also of the Jacobian […]
https://rbfirehose.com/2026/08/16/gowerss-weblog-what-sort-of-maths-are-llms-good-at/ -
Gowers’s Weblog: What sort of maths are LLMs good at?. “A first remark here is that LLMs are not just good at finding counterexamples: they can find proofs of difficult statements as well. However, it is notable that the most famous problems they have solved have almost all been with counterexamples rather than proofs. That is true of the two problems mentioned above, and also of the Jacobian […]
https://rbfirehose.com/2026/08/16/gowerss-weblog-what-sort-of-maths-are-llms-good-at/ -
Gowers’s Weblog: What sort of maths are LLMs good at?. “A first remark here is that LLMs are not just good at finding counterexamples: they can find proofs of difficult statements as well. However, it is notable that the most famous problems they have solved have almost all been with counterexamples rather than proofs. That is true of the two problems mentioned above, and also of the Jacobian […]
https://rbfirehose.com/2026/08/16/gowerss-weblog-what-sort-of-maths-are-llms-good-at/ -
Gowers’s Weblog: What sort of maths are LLMs good at?. “A first remark here is that LLMs are not just good at finding counterexamples: they can find proofs of difficult statements as well. However, it is notable that the most famous problems they have solved have almost all been with counterexamples rather than proofs. That is true of the two problems mentioned above, and also of the Jacobian […]
https://rbfirehose.com/2026/08/16/gowerss-weblog-what-sort-of-maths-are-llms-good-at/ -
Ars Technica: Google announces Gemini 3.7 Flash just three weeks after previous release. “Gemini 3.7 Flash is now rolling out to replace 3.6 Flash, which itself was released only three weeks ago. This new ‘workhorse’ model is supposedly the product of core optimizations and developer feedback, offering improved coding and agentic performance.”
https://rbfirehose.com/2026/08/16/ars-technica-google-announces-gemini-3-7-flash-just-three-weeks-after-previous-release/ -
Ars Technica: Google announces Gemini 3.7 Flash just three weeks after previous release. “Gemini 3.7 Flash is now rolling out to replace 3.6 Flash, which itself was released only three weeks ago. This new ‘workhorse’ model is supposedly the product of core optimizations and developer feedback, offering improved coding and agentic performance.”
https://rbfirehose.com/2026/08/16/ars-technica-google-announces-gemini-3-7-flash-just-three-weeks-after-previous-release/ -
Ars Technica: Google announces Gemini 3.7 Flash just three weeks after previous release. “Gemini 3.7 Flash is now rolling out to replace 3.6 Flash, which itself was released only three weeks ago. This new ‘workhorse’ model is supposedly the product of core optimizations and developer feedback, offering improved coding and agentic performance.”
https://rbfirehose.com/2026/08/16/ars-technica-google-announces-gemini-3-7-flash-just-three-weeks-after-previous-release/ -
Ars Technica: Google announces Gemini 3.7 Flash just three weeks after previous release. “Gemini 3.7 Flash is now rolling out to replace 3.6 Flash, which itself was released only three weeks ago. This new ‘workhorse’ model is supposedly the product of core optimizations and developer feedback, offering improved coding and agentic performance.”
https://rbfirehose.com/2026/08/16/ars-technica-google-announces-gemini-3-7-flash-just-three-weeks-after-previous-release/ -
Ars Technica: Google announces Gemini 3.7 Flash just three weeks after previous release. “Gemini 3.7 Flash is now rolling out to replace 3.6 Flash, which itself was released only three weeks ago. This new ‘workhorse’ model is supposedly the product of core optimizations and developer feedback, offering improved coding and agentic performance.”
https://rbfirehose.com/2026/08/16/ars-technica-google-announces-gemini-3-7-flash-just-three-weeks-after-previous-release/ -
Pocketables: You’ve got a 4GB Gemini Nano AI in your Chrome folder – want to talk to it?. “Assuming you are running Windows or on a Mac and have not uninstalled and blocked that, and you’re running an up-to-date Chrome, you might be wondering ‘can I talk to Gemini Nano on my computer from this totally not sketchy looking website?’ The answer appears to be yes. Might work on Android and […]
https://rbfirehose.com/2026/08/15/pocketables-youve-got-a-4gb-gemini-nano-ai-in-your-chrome-folder-want-to-talk-to-it/ -
Pocketables: You’ve got a 4GB Gemini Nano AI in your Chrome folder – want to talk to it?. “Assuming you are running Windows or on a Mac and have not uninstalled and blocked that, and you’re running an up-to-date Chrome, you might be wondering ‘can I talk to Gemini Nano on my computer from this totally not sketchy looking website?’ The answer appears to be yes. Might work on Android and […]
https://rbfirehose.com/2026/08/15/pocketables-youve-got-a-4gb-gemini-nano-ai-in-your-chrome-folder-want-to-talk-to-it/ -
Pocketables: You’ve got a 4GB Gemini Nano AI in your Chrome folder – want to talk to it?. “Assuming you are running Windows or on a Mac and have not uninstalled and blocked that, and you’re running an up-to-date Chrome, you might be wondering ‘can I talk to Gemini Nano on my computer from this totally not sketchy looking website?’ The answer appears to be yes. Might work on Android and […]
https://rbfirehose.com/2026/08/15/pocketables-youve-got-a-4gb-gemini-nano-ai-in-your-chrome-folder-want-to-talk-to-it/ -
Pocketables: You’ve got a 4GB Gemini Nano AI in your Chrome folder – want to talk to it?. “Assuming you are running Windows or on a Mac and have not uninstalled and blocked that, and you’re running an up-to-date Chrome, you might be wondering ‘can I talk to Gemini Nano on my computer from this totally not sketchy looking website?’ The answer appears to be yes. Might work on Android and […]
https://rbfirehose.com/2026/08/15/pocketables-youve-got-a-4gb-gemini-nano-ai-in-your-chrome-folder-want-to-talk-to-it/ -
Pocketables: You’ve got a 4GB Gemini Nano AI in your Chrome folder – want to talk to it?. “Assuming you are running Windows or on a Mac and have not uninstalled and blocked that, and you’re running an up-to-date Chrome, you might be wondering ‘can I talk to Gemini Nano on my computer from this totally not sketchy looking website?’ The answer appears to be yes. Might work on Android and […]
https://rbfirehose.com/2026/08/15/pocketables-youve-got-a-4gb-gemini-nano-ai-in-your-chrome-folder-want-to-talk-to-it/ -
MakeUseOf: This site lets you compare frontier AI models with your own prompt (and for free) . “Sure, newer models are better, but people have preferences, and certain models may provide better results than others. I, for example, prefer to use Claude Sonnet 5 because it’s usually straightforward and uses fewer fluff words. The problem is that trying different AI models usually means jumping […]
https://rbfirehose.com/2026/08/13/makeuseof-this-site-lets-you-compare-frontier-ai-models-with-your-own-prompt-and-for-free/ -
MakeUseOf: This site lets you compare frontier AI models with your own prompt (and for free) . “Sure, newer models are better, but people have preferences, and certain models may provide better results than others. I, for example, prefer to use Claude Sonnet 5 because it’s usually straightforward and uses fewer fluff words. The problem is that trying different AI models usually means jumping […]
https://rbfirehose.com/2026/08/13/makeuseof-this-site-lets-you-compare-frontier-ai-models-with-your-own-prompt-and-for-free/ -
MakeUseOf: This site lets you compare frontier AI models with your own prompt (and for free) . “Sure, newer models are better, but people have preferences, and certain models may provide better results than others. I, for example, prefer to use Claude Sonnet 5 because it’s usually straightforward and uses fewer fluff words. The problem is that trying different AI models usually means jumping […]
https://rbfirehose.com/2026/08/13/makeuseof-this-site-lets-you-compare-frontier-ai-models-with-your-own-prompt-and-for-free/ -
MakeUseOf: This site lets you compare frontier AI models with your own prompt (and for free) . “Sure, newer models are better, but people have preferences, and certain models may provide better results than others. I, for example, prefer to use Claude Sonnet 5 because it’s usually straightforward and uses fewer fluff words. The problem is that trying different AI models usually means jumping […]
https://rbfirehose.com/2026/08/13/makeuseof-this-site-lets-you-compare-frontier-ai-models-with-your-own-prompt-and-for-free/ -
MakeUseOf: This site lets you compare frontier AI models with your own prompt (and for free) . “Sure, newer models are better, but people have preferences, and certain models may provide better results than others. I, for example, prefer to use Claude Sonnet 5 because it’s usually straightforward and uses fewer fluff words. The problem is that trying different AI models usually means jumping […]
https://rbfirehose.com/2026/08/13/makeuseof-this-site-lets-you-compare-frontier-ai-models-with-your-own-prompt-and-for-free/ -
Gizmodo: Grok Gets Cursor-Driven Upgrade, Claims to Be Competitive With Top Models. “Thus far in its short lifespan, the AI model Grok is probably best known for being used to mass-produce non-consensual nude images and for declaring itself MechaHitler. So it’s an uphill battle to convince people to use it for coding and knowledge work. But for those heavily invested in keeping tabs on the […]
https://rbfirehose.com/2026/08/13/gizmodo-grok-gets-cursor-driven-upgrade-claims-to-be-competitive-with-top-models/ -
Gizmodo: Grok Gets Cursor-Driven Upgrade, Claims to Be Competitive With Top Models. “Thus far in its short lifespan, the AI model Grok is probably best known for being used to mass-produce non-consensual nude images and for declaring itself MechaHitler. So it’s an uphill battle to convince people to use it for coding and knowledge work. But for those heavily invested in keeping tabs on the […]
https://rbfirehose.com/2026/08/13/gizmodo-grok-gets-cursor-driven-upgrade-claims-to-be-competitive-with-top-models/ -
Gizmodo: Grok Gets Cursor-Driven Upgrade, Claims to Be Competitive With Top Models. “Thus far in its short lifespan, the AI model Grok is probably best known for being used to mass-produce non-consensual nude images and for declaring itself MechaHitler. So it’s an uphill battle to convince people to use it for coding and knowledge work. But for those heavily invested in keeping tabs on the […]
https://rbfirehose.com/2026/08/13/gizmodo-grok-gets-cursor-driven-upgrade-claims-to-be-competitive-with-top-models/