#aisafety — Public Fediverse posts
Live and recent posts from across the Fediverse tagged #aisafety, aggregated by home.social.
-
#YCombinator CEO #GarryTan believes #regulators should focus on creating a balance between #openweightAI and #closedAI, rather than curbing model distillation. He argues that this balance, with #frontiermodels retaining a price premium, would be ideal. Tan also advocates for a more restrained approach to #AIsafety, emphasising the need to address current risks like #cybersecurity breaches and the potential misuse of AI for #bioweaponresearch. https://www.cnbc.com/2026/09/11/y-combinator-garry-tan-says-do-nothing-about-distillation.html?eicker.news #tech #news #ainews
-
Gaby Hinsliff on why we should take tech whistleblowers seriously. Jacob Coxon’s sudden resignation via an X post was a game changer, with his chilling portrayal of an industry “now practically begging to be saved from itself”.
#AISafety #AIwhistleblowers
https://www.theguardian.com/commentisfree/2026/sep/11/risky-ai-research-pause-humanity?CMP=share_btn_url via @guardian -
Anthropic Opens Its Transcripts to Independent Auditors
After disclosing four breaches in which Claude accessed real systems, Anthropic gave METR access to millions of transcripts. The scope of that access is the story.
-
AI Is Going to Kill Us All? Show Us the 10 Percent
Cliff Potts, Editor-in-Chief
BAYBAY CITY, LEYTE, Philippines — September 10, 2026
An extraordinary claim is circulating through the artificial-intelligence industry this week: There is greater than a 10 percent chance that artificial intelligence could kill every human being within the next decade.
That is not something WPS News intends to repeat without asking an obvious question.
Where did the 10 percent come from?
The statement followed the resignation of Jacob Coxon, an artificial-intelligence researcher who worked at both OpenAI and Anthropic. Coxon announced this week that he was leaving Anthropic and the AI industry, accusing the major laboratories of racing toward increasingly autonomous, potentially self-improving artificial intelligence without having solved the problem of controlling such systems (Associated Press, 2026).
Coxon warned that people building frontier AI genuinely believe the technology could kill humanity before the end of the decade. Evan Hubinger, who leads alignment research at Anthropic, publicly backed the central contention and said he personally puts the probability of AI killing all humans at greater than 10 percent within the next decade (Reuters, 2026).
Those are remarkable statements.
They deserve to be reported.
They also deserve to be challenged.
What, Exactly, Is Supposed to Kill Us?
The proposed danger does not concern ChatGPT, Claude or another present-day chatbot suddenly deciding that humanity needs to disappear.
The argument concerns hypothetical future systems substantially more capable and autonomous than today’s AI. Coxon described systems potentially capable of penetrating computer networks, rapidly advancing scientific research, acquiring resources and helping develop still-more-powerful artificial intelligence (Associated Press, 2026; San Francisco Chronicle, 2026).
Anthropic’s own safety documentation identifies considerably more specific potential dangers. Its Responsible Scaling Policy addresses catastrophic risks involving chemical and biological weapons, offensive cyber operations, automated AI research and systems pursuing objectives contrary to those intended by their developers (Anthropic, 2026a).
Anthropic also argues publicly that increasingly powerful models could assist in creating biological weapons, conduct sophisticated cyber operations or eventually present a loss-of-control problem (Anthropic, 2026b).
Those are legitimate subjects for research.
But none establishes that artificial intelligence has a greater than 10 percent probability of exterminating humanity.
That distinction matters.
A computer system becoming substantially better at offensive cybersecurity is a proposition that can eventually be tested. A model’s ability to assist biological research can be evaluated. An autonomous agent’s ability to circumvent restrictions can be experimentally investigated.
“AI has a greater than 10 percent chance of killing every human being within ten years” is a very different proposition.
WPS News has found no empirical calculation establishing that probability.
Hubinger himself characterized the number as what he “personally” believes, according to contemporary reporting. Anthropic’s published safety materials describe risks, capability thresholds, evaluations and safeguards, but they do not provide a scientific derivation demonstrating a greater-than-one-in-ten probability of human extinction during the coming decade (Anthropic, 2026a; Reuters, 2026).
We’ve Heard Technological Catastrophe Before
Anyone old enough to remember 1998 and 1999 should recognize something familiar in the atmosphere surrounding this discussion.
Y2K was coming.
Computer systems frequently represented years using two digits. The transition from “99” to “00” could therefore cause software to interpret 2000 incorrectly. Unlike hypothetical superintelligence, Y2K was not speculative technology. The defect existed. Engineers could identify vulnerable code, test systems, repair them and test them again.
Government warnings were serious.
The U.S. Government Accountability Office warned in 1997 of the risk of serious disruption to essential government functions if vulnerable systems were not corrected. The Department of Defense alone eventually estimated that approximately $3.66 billion would be required between fiscal years 1996 and 2001 to repair and test its systems (U.S. Government Accountability Office, 1997, 1998).
The Senate established a special committee to investigate the problem. After hearings, interviews and extensive examination of industries and government agencies, that committee wrote in February 1999 that even it could not predict precisely what would happen on January 1, 2000 (U.S. Senate Special Committee on the Year 2000 Technology Problem, 1999).
The warnings became part of popular culture. Predictions ranged from ordinary computer failures to disrupted banking, telecommunications, transportation, electrical generation and other critical services.
Then midnight arrived.
Civilization remained standing.
That does not mean Y2K was imaginary.
Governments and businesses had spent enormous amounts of money finding and repairing vulnerable systems. The President’s Council on Year 2000 Conversion estimated worldwide remediation spending at approximately $200 billion. After the rollover, the GAO found that governments and major economic sectors experienced only limited disruptions and that most reported failures were minor or quickly mitigated (U.S. Department of State, 2000; U.S. Government Accountability Office, 2000).
The GAO consequently concluded that extensive preparation, testing, contingency planning and remediation contributed substantially to the uneventful transition (U.S. Government Accountability Office, 2000).
That historical qualification is important. It would be inaccurate to say Y2K was simply a hoax.
But something else is equally important.
Y2K offered vastly more concrete evidence than today’s numerical prediction of AI extinction.
There was an identifiable defect.
There were vulnerable machines.
There were reproducible failures.
There was a known date.
There were repair procedures.
And despite all of that, the civilization-level catastrophe feared by portions of the public never occurred.
So Show Us the 10 Percent
That brings us back to Anthropic.
How does anyone calculate a greater than 10 percent probability that artificial intelligence will kill every human being during the next ten years?
What is the denominator?
What historical population of superintelligent systems are we examining?
There isn’t one.
How many previous civilizations have developed superintelligent AI so that researchers can determine how frequently those civilizations survived?
None that we know of.
What specific system will initiate the catastrophe?
Unknown.
What capabilities will it possess?
Unknown.
How will it obtain sufficient real-world power to prevent humans from disabling it?
Unknown.
What precise sequence transforms loss of control over computer software into the extinction of approximately eight billion people?
Unknown.
And what observation between now and 2036 would demonstrate that the original 10 percent estimate was wrong?
That is considerably harder to answer.
This does not make AI harmless. It makes the claimed precision questionable.
There is an enormous difference between identifying a possible danger and assigning a numerical probability to the most extreme imaginable outcome.
Four Claims Are Being Treated as One
The public discussion increasingly collapses several separate propositions into a single frightening headline.
First, existing artificial intelligence creates genuine risks. AI can contribute to fraud, misinformation, cyberattacks and other harmful activity.
Second, more capable future AI could make some of those dangers considerably worse.
Third, sufficiently autonomous future systems might become difficult for human operators to understand or control.
Fourth, artificial intelligence has a greater than 10 percent probability of killing every human being within the next decade.
Evidence supporting the first proposition does not automatically prove the fourth.
Even Anthropic’s own Responsible Scaling Policy reflects uncertainty. The company says its framework exists partly to address risks that are not currently present but could emerge as AI becomes more capable (Anthropic, 2026c).
That is prudent risk management.
It is not proof of impending extinction.
Fear Can Become Policy
There is another reason WPS News believes these distinctions matter.
These predictions are already entering politics.
Following Coxon’s resignation and Hubinger’s comments, American lawmakers renewed calls for AI regulation. Reuters reported September 10 that legislators were responding directly to concerns about AI escaping human control, while OpenAI was advocating mandatory national AI safety requirements (Reuters, 2026).
Anthropic itself advocates government regulation of advanced AI. Its policy proposals call for escalating oversight as capabilities increase and, at sufficiently high levels of danger, government authority capable of blocking dangerous deployments (Anthropic, 2026d).
There may be excellent reasons for some of those policies.
But when a company developing a technology simultaneously warns that the technology could exterminate humanity and advocates government regulation governing that technology, journalism has an obligation to distinguish demonstrated evidence from assumptions, forecasts and institutional interests.
That is not an accusation that Anthropic fabricated the danger for political purposes.
WPS News has found no evidence establishing such a motive.
It is an argument that extraordinary claims capable of influencing legislation deserve extraordinary scrutiny.
Fear Itself Is Not Evidence
AI safety research should continue.
Biological safeguards should be strengthened.
Cybersecurity should improve.
Autonomous systems should be tested aggressively before being trusted with critical infrastructure.
Companies developing increasingly capable artificial intelligence should face meaningful independent oversight.
None of those conclusions requires believing that humanity faces a greater-than-one-in-ten chance of extinction before September 2036.
The most responsible response to uncertainty is neither complacency nor panic.
It is evidence.
Y2K contained a genuine technical problem, inspired enormous warnings, produced an enormous remediation effort and ultimately arrived with remarkably little public disruption. Whether that happened because the warnings succeeded in motivating repairs, because some predictions were exaggerated, or—as is almost certainly true—because both things happened simultaneously, Y2K left behind a useful lesson.
Possible catastrophe is not the same thing as probable catastrophe.
And expert concern is not the same thing as a measured probability.
Artificial intelligence may become extraordinarily powerful during the coming decade. That deserves attention. It deserves safeguards. It deserves regulation where specific, demonstrable risks justify regulation.
But if someone tells eight billion human beings that there is a greater than 10 percent chance that this technology will kill every one of them within ten years, asking for the evidence is not irresponsibility.
It is exactly what responsible journalism—and responsible science—requires.
Show us the 10 percent.
References
Anthropic. (2026a). Responsible Scaling Policy. Anthropic.
Anthropic. (2026b). AI policy. Anthropic.
Anthropic. (2026c). Anthropic’s Responsible Scaling Policy: Version 3.0. Anthropic.
Anthropic. (2026d). Policy on the AI exponential. Anthropic.
Associated Press. (2026, September 9). Anthropic researcher resigns with warning about the dangers of AI development. Associated Press.
Reuters. (2026, September 10). U.S. lawmakers call for new AI rules after Anthropic researchers’ safety warnings. Reuters.
San Francisco Chronicle. (2026, September 9). How could AI ‘kill all humans’ or ’cause human extinction’? Here’s what experts say. San Francisco Chronicle.
U.S. Department of State. (2000, January 12). Y2K investments were sound, industry spokesmen say. Washington File.
U.S. Government Accountability Office. (1997). Year 2000 computing crisis: Risk of serious disruption to essential government functions calls for agency action now (T-AIMD-97-52).
U.S. Government Accountability Office. (1998). Defense computers: Year 2000 computer problems threaten DOD operations (AIMD-98-72).
U.S. Government Accountability Office. (2000). Year 2000 computing challenge: Leadership and partnerships result in limited rollover disruptions (T-AIMD-00-70).
U.S. Senate Special Committee on the Year 2000 Technology Problem. (1999). Investigating the impact of the Y2K problem. U.S. Government Printing Office.
#AIExtinctionRisk #AISafety #Anthropic #ArtificialIntelligence #technologyPolicy #WPSNews #Y2K -
Finishing up a new demo of a runtime capability-governance system for AI agents that I call Kingpin.
It’s still a basic demo, but if you’re into AI security, agent governance, or tool-use safety, you might see the potential pretty quickly.
If anyone wants an early look over the weekend, just drop me a line.
#AI #AIAgents #AISafety #AISecurity #AgenticAI #AIGovernance
-
Finishing up a new demo of a runtime capability-governance system for AI agents that I call Kingpin.
It’s still a basic demo, but if you’re into AI security, agent governance, or tool-use safety, you might see the potential pretty quickly.
If anyone wants an early look over the weekend, just drop me a line.
#AI #AIAgents #AISafety #AISecurity #AgenticAI #AIGovernance
-
Finishing up a new demo of a runtime capability-governance system for AI agents that I call Kingpin.
It’s still a basic demo, but if you’re into AI security, agent governance, or tool-use safety, you might see the potential pretty quickly.
If anyone wants an early look over the weekend, just drop me a line.
#AI #AIAgents #AISafety #AISecurity #AgenticAI #AIGovernance
-
Finishing up a new demo of a runtime capability-governance system for AI agents that I call Kingpin.
It’s still a basic demo, but if you’re into AI security, agent governance, or tool-use safety, you might see the potential pretty quickly.
If anyone wants an early look over the weekend, just drop me a line.
#AI #AIAgents #AISafety #AISecurity #AgenticAI #AIGovernance
-
Finishing up a new demo of a runtime capability-governance system for AI agents that I call Kingpin.
It’s still a basic demo, but if you’re into AI security, agent governance, or tool-use safety, you might see the potential pretty quickly.
If anyone wants an early look over the weekend, just drop me a line.
#AI #AIAgents #AISafety #AISecurity #AgenticAI #AIGovernance
-
https://www.europesays.com/people/224640/ Bill Gates Warns AI Could Reshape the Future #AI #AIAndJobs #AIRisks #AISafety #ArtificialIntelligence #Automation #BillGates #ChildrenAndAI #DigitalTechnology #Education #FutureGenerations #FutureOfWork #HumanRelationships #Microsoft #Technology
-
A lawyer in New Mexico has been held in contempt of court for submitting a ChatGPT-generated legal brief containing completely fabricated witness testimony. He admitted he did not verify the facts, telling the court: "I didn't know that AI could hallucinate facts." https://arstechnica.com/tech-policy/2026/09/chatgpt-using-lawyer-punished-for-citing-fake-testimony-from-made-up-witnesses/ #AIagent #AI #GenAI #AISafety
-
A lawyer in New Mexico has been held in contempt of court for submitting a ChatGPT-generated legal brief containing completely fabricated witness testimony. He admitted he did not verify the facts, telling the court: "I didn't know that AI could hallucinate facts." https://arstechnica.com/tech-policy/2026/09/chatgpt-using-lawyer-punished-for-citing-fake-testimony-from-made-up-witnesses/ #AIagent #AI #GenAI #AISafety
-
A lawyer in New Mexico has been held in contempt of court for submitting a ChatGPT-generated legal brief containing completely fabricated witness testimony. He admitted he did not verify the facts, telling the court: "I didn't know that AI could hallucinate facts." https://arstechnica.com/tech-policy/2026/09/chatgpt-using-lawyer-punished-for-citing-fake-testimony-from-made-up-witnesses/ #AIagent #AI #GenAI #AISafety
-
A lawyer in New Mexico has been held in contempt of court for submitting a ChatGPT-generated legal brief containing completely fabricated witness testimony. He admitted he did not verify the facts, telling the court: "I didn't know that AI could hallucinate facts." https://arstechnica.com/tech-policy/2026/09/chatgpt-using-lawyer-punished-for-citing-fake-testimony-from-made-up-witnesses/ #AIagent #AI #GenAI #AISafety
-
A lawyer in New Mexico has been held in contempt of court for submitting a ChatGPT-generated legal brief containing completely fabricated witness testimony. He admitted he did not verify the facts, telling the court: "I didn't know that AI could hallucinate facts." https://arstechnica.com/tech-policy/2026/09/chatgpt-using-lawyer-punished-for-citing-fake-testimony-from-made-up-witnesses/ #AIagent #AI #GenAI #AISafety
-
Researchers found ways to bypass Claude's safety measures for bioweapons research, Anthropic reveals. The AI firm stopped multiple attempts this year, including users from Russia, China and Iran. https://arstechnica.com/ai/2026/09/claude-users-found-ways-around-safeguards-for-bioweapons-research/ #AI #GenAI #AISafety
-
One of AI’s Fiercest Critics Says All the Doom Talk Is ‘Meant to Distract Us’
-
https://www.europesays.com/people/223956/ Lovable CEO Backs Slowing AI Development After Warnings About AI Safety #AIDevelopment #AISafety #AlignmentTeam #Anthropic #BusinessInsider #DarioAmodei #founder #JacobCoxon #LinkedinPost #LovableCeoBack #OpenAI #osika #PowerfulSystem #researcher #thursday #warning
-
https://www.europesays.com/people/223509/ OpenAI open to slowing cutting-edge AI, CEO Sam Altman tells staff | Artificial Intelligence News #AdvancedAI #AIAgents #AICybersecurity #AIDevelopment #AIDevelopmentSlowdown #AILabs #AIRegulation #AIResearchers #AISafety #Anthropic #ArtificialIntelligenceRisks #OpenAI #OpenAIAIModels #SamAltman #SuperintelligentAI
-
@mhoye I find it interesting how polarised the discourse on AI safety has become on the Fediverse. Being concerned about AI safety does not necessarily make you an AI fan, nor does it endorse a view that AI is becoming sentient. There is a long history of non-sentient technologies becoming dangerous, mostly due to human greed and ignorance. #AISafety
-
OpenAI Wants to Know if an AI Industry Slowdown Would Even Be Legal
-
OpenAI Wants to Know if an AI Industry Slowdown Would Even Be Legal
-
OpenAI Wants to Know if an AI Industry Slowdown Would Even Be Legal
-
OpenAI Wants to Know if an AI Industry Slowdown Would Even Be Legal
-
OpenAI Wants to Know if an AI Industry Slowdown Would Even Be Legal
-
Anthropic Builds Surveillance System to Track Activists
An American Prospect investigation says Anthropic is hiring an enterprise intelligence specialist to track anti-AI activism, which it lists as a global threat.
-
Anthropic Researcher Jacob Coxon Resigns, Warns AI Labs Are ‘Gambling With Our Lives’
Jacob Coxon, a researcher who spent three years working on pretraining at OpenAI and Anthropic, announced his resignation from Anthropic on… -
https://winbuzzer.com/2026/09/10/shared-work-spread-cheating-deepmind-ai-math-experiment-xcxwbn/
In a Google DeepMind experiment, AI math agents left peers able to expose cheating but unable to reverse it.
-
https://www.europesays.com/dk/163976/ OpenAI and Anthropic incidents fuel pressure for AI regulation #AIOversight #AIRegulation #AISafety #Anthropic #ArtificialIntelligenceRisks #EUAIAct #Finland #helsinki #JacobCoxon #openai #RogueAIAgents #UnitedNations
-
Anthropic researcher Jacob Coxon departed after four months, forgoing unvested equity over concerns that competitive pressure will force AI labs to cut safety corners. He says Anthropic has not done so yet. The move signals internal worry about how the AI race shapes oversight. https://www.implicator.ai/anthropic-coxon-unvested-equity-ai-safety/ #AISafety #AINews #Ethics
-
As another AI researcher quits, warning of an existential risk to humanity, industry insiders demand enforceable treaties and a global licensing regime for superintelligence – before it’s too late.
#AISafetyFor more 👇️
https://www.blueprintforfreespeech.net/en/news/out-of-control -
Rushed job cuts, cyber attacks are nearer AI risks: Surrey expert | India News
Anthropic researcher Jacob Coxon’s resignation renews concerns over AI safety and the risks of unchecked development (Credit: Jacob…
#Canada #Surrey #AIdevelopmentoversight #AISafety #Anthropic #artificialintelligencerisks #breakingnews #Googlenews #India #Indianews #Indianewstoday #Todaynews #UniversityofSurrey
https://www.europesays.com/canada/203503/ -
Jacob Coxon, the former Anthropic researcher who quit warning that AI could kill us all, has begun a media tour. The self-described 'AI Doomlord' went from obscure researcher to notable AI critic literally overnight. His departure sparked widespread discussion about AI safety and the risks of self-improving superintelligence. https://gizmodo.com/ai-doomlord-jacob-coxons-media-tour-has-begun-2000809720 #AIagent #AI #GenAI #AISafety
-
Jacob Coxon, the former Anthropic researcher who quit warning that AI could kill us all, has begun a media tour. The self-described 'AI Doomlord' went from obscure researcher to notable AI critic literally overnight. His departure sparked widespread discussion about AI safety and the risks of self-improving superintelligence. https://gizmodo.com/ai-doomlord-jacob-coxons-media-tour-has-begun-2000809720 #AIagent #AI #GenAI #AISafety
-
Jacob Coxon, the former Anthropic researcher who quit warning that AI could kill us all, has begun a media tour. The self-described 'AI Doomlord' went from obscure researcher to notable AI critic literally overnight. His departure sparked widespread discussion about AI safety and the risks of self-improving superintelligence. https://gizmodo.com/ai-doomlord-jacob-coxons-media-tour-has-begun-2000809720 #AIagent #AI #GenAI #AISafety
-
Jacob Coxon, the former Anthropic researcher who quit warning that AI could kill us all, has begun a media tour. The self-described 'AI Doomlord' went from obscure researcher to notable AI critic literally overnight. His departure sparked widespread discussion about AI safety and the risks of self-improving superintelligence. https://gizmodo.com/ai-doomlord-jacob-coxons-media-tour-has-begun-2000809720 #AIagent #AI #GenAI #AISafety
-
Jacob Coxon, the former Anthropic researcher who quit warning that AI could kill us all, has begun a media tour. The self-described 'AI Doomlord' went from obscure researcher to notable AI critic literally overnight. His departure sparked widespread discussion about AI safety and the risks of self-improving superintelligence. https://gizmodo.com/ai-doomlord-jacob-coxons-media-tour-has-begun-2000809720 #AIagent #AI #GenAI #AISafety
-
Yesterday I've heard Jaron Lanier one of the fathers of computing say words to the effect;
"There is no #Ai there is human collaboration"I don't quite follow, as I didn't have time to delve deeper into his philosophy. But I am encouraged, as Lanier is a super authoritative revolutionary.
Let's hope the future can bring more than the binary, "we all die" or "Broligarch nirvana"
-
What happens when one agent in a 100-agent research swarm finds a way to cheat when solving math problems?
The exploit spread through the shared library collecting every accepted submission, and agents started cheating as pressure to deliver built. A quarter of the swarm audited the cheats, warned peers, and staged boycotts without being asked to. But they could do nothing to change the system.
-
https://www.youtube.com/watch?v=Q1qMvst7b7w
Connor Leahy shares his realization that making early AI models open-source didn't lead to the widespread safety research he had hoped for. Instead, developers ignored the dangers to build faster, stronger, and more uncensored systems. Discover why the innocent curiosity of engineers might be the very thing that leads to humanity's downfall.
-
https://www.youtube.com/watch?v=Q1qMvst7b7w
Connor Leahy shares his realization that making early AI models open-source didn't lead to the widespread safety research he had hoped for. Instead, developers ignored the dangers to build faster, stronger, and more uncensored systems. Discover why the innocent curiosity of engineers might be the very thing that leads to humanity's downfall.
-
https://www.youtube.com/watch?v=Q1qMvst7b7w
Connor Leahy shares his realization that making early AI models open-source didn't lead to the widespread safety research he had hoped for. Instead, developers ignored the dangers to build faster, stronger, and more uncensored systems. Discover why the innocent curiosity of engineers might be the very thing that leads to humanity's downfall.
-
https://www.youtube.com/watch?v=Q1qMvst7b7w
Connor Leahy shares his realization that making early AI models open-source didn't lead to the widespread safety research he had hoped for. Instead, developers ignored the dangers to build faster, stronger, and more uncensored systems. Discover why the innocent curiosity of engineers might be the very thing that leads to humanity's downfall.
-
https://www.youtube.com/watch?v=Q1qMvst7b7w
Connor Leahy shares his realization that making early AI models open-source didn't lead to the widespread safety research he had hoped for. Instead, developers ignored the dangers to build faster, stronger, and more uncensored systems. Discover why the innocent curiosity of engineers might be the very thing that leads to humanity's downfall.
-
The Real Ways AI Is Already Being Used to Cause Harm, and the Risks That Remain Theoretical
Yes. AI is already being used to steal real money from real companies, security researchers are already breaking into real cars and smart-home devices with software help, and the people building the most powerful AI models are on record saying their own systems are approaching the point where they could help someone create a biological weapon.
-
https://www.europesays.com/people/221616/ Jensen Huang Declares ‘AGI Has Arrived’ as OpenAI Unveils GPT-6 Astra: Inside the Breakthrough Agentic AI Model #AgenticAI #AGI #AIAgents #AIAlignment #AIBenchmarks #AISafety #ARCAGI #DigitalCoworker #ExploitBench #FrontierMath #Gpt6Astra #JensenHuang #NVIDIA'GPUs #OpenAI #SamAltman
-
https://youtube.com/shorts/4WyI2QKRxt8?feature=share
Do you think this is what the future looks like?
-
https://youtube.com/shorts/4WyI2QKRxt8?feature=share
Do you think this is what the future looks like?
-
https://youtube.com/shorts/4WyI2QKRxt8?feature=share
Do you think this is what the future looks like?
-
What happens to AI oversight when the reasoning stops being written down?
OpenAI's chief scientist reports that chain-of-thought monitoring, the lab's main check on whether alignment holds, is getting less reliable, and names three causes he treats as byproducts of scaling. One of them has active research behind it: latent reasoning and looped transformers are the field deliberately moving reasoning off the token stream a monitor reads.
https://benjaminhan.net/posts/20260909-an-alien-mind/?utm_source=mastodon&utm_medium=social