#aisafety — Public Fediverse posts
Live and recent posts from across the Fediverse tagged #aisafety, aggregated by home.social.
-
OpenAI is testing safety signals that detect misuse across related interactions without retaining customer prompts or responses. The key question is auditability: can customers reproduce, investigate, and appeal an alert when the provider holds only the signal—not the evidence? https://www.computerworld.com/article/4212412/openai-adds-an-ai-safety-layer-to-detect-misuse-without-retaining-enterprise-data-2.html #AISafety #Privacy
-
How Do We Know Whether AI Is Actually Helping People?
What several AI models said when we asked them the same question
Artificial intelligence is getting more capable very quickly. It can write, analyze data, create images, translate languages, help with research, and solve problems that once required trained specialists.
But greater capability does not automatically mean a better life for people.
That was the starting point for a small cross-model experiment. We asked several AI systems the same basic question:
How would you determine whether increasingly capable AI is actually benefiting human life?
We also invited each model to question the premise, redefine the problem, or suggest something better than a single index. The models were instructed to answer independently without browsing the web or using outside tools.
The responses differed in style and emphasis. Some focused on measurable outcomes. Others focused on human dignity, democratic participation, meaningful work, or the danger of becoming dependent on systems we do not control.
Yet a surprisingly clear agreement emerged.
Capability is not the same as benefit
Technical progress is easy to measure. We can count how many problems an AI solves, how quickly it works, or how well it performs on tests.
Human flourishing is harder to measure. It includes health, safety, freedom, relationships, purpose, knowledge, creativity, and the ability to shape one’s own life.
An AI system may become better at achieving a goal while the goal itself harms people. A highly effective system might increase surveillance, spread convincing scams, replace human judgment, concentrate power, or keep users engaged at the expense of their attention and well-being.
So the important question is not simply, “What can AI do?”
It is:
What becomes possible for people because of AI—and what becomes more difficult, fragile, or impossible?
Look at human outcomes, not just machine performance
Across the responses, the models repeatedly shifted attention away from the machine and toward human life.
They suggested asking whether people are:
- healthier and safer;
- more financially secure;
- better able to learn and create;
- more connected to other people;
- more informed without being manipulated;
- able to understand and challenge important decisions;
- free to refuse the technology or choose another path.
This also requires examining harms, not merely counting success stories. Time saved by one group may come with unemployment, stress, lost privacy, or reduced opportunity for another.
A true evaluation must ask who receives the benefits, who carries the risks, and who has the power to decide.
Agency belongs at the center
One of the strongest shared themes was human agency: our ability to understand, choose, refuse, act, and take responsibility.
Convenience alone is not agency. A system can make life easier while quietly reducing a person’s choices or replacing their judgment.
Helpful AI should strengthen people’s ability to participate in their own lives. It should make important decisions more understandable, provide meaningful options, and allow people to correct mistakes or appeal harmful outcomes.
People need more than access to AI. They need power in relation to it.
Assistance should not erase human competence
Several responses warned that a tool can help us today while making us less capable tomorrow.
If people lose the knowledge needed to check an AI system, operate without it, or recover when it fails, short-term convenience may create long-term fragility.
This suggests a simple test:
If the AI disappeared tomorrow, what knowledge, skill, judgment, and institutional capacity would remain?
The best systems may act more like scaffolding than substitutes. Scaffolding helps people reach farther while they continue developing their own abilities. Substitution can slowly remove the very competence that makes human oversight possible.
Benefit is not one number
Another broad agreement was that a single “AI Benefit Score” would hide too much.
An average can make widespread gains look impressive while concealing serious harm to a smaller or less powerful group. One number can also allow gains in productivity to cancel out losses of privacy, dignity, freedom, or democratic control.
A better approach would combine several forms of evaluation:
- Outcomes: Are people healthier, safer, more secure, more connected, and materially better off?
- Agency: Are people more able to choose, understand, refuse, create, and govern their lives?
- Resilience: Are human skills, social institutions, alternatives, and the ability to recover being preserved?
Each of these should be examined across four additional questions:
- Distribution: Who benefits, and who is harmed?
- Power: Who controls the system and can be held accountable?
- Time: What happens months, years, or generations later?
- Causation: Did AI actually cause the change, or did it merely appear alongside it?
Some harms may also require firm boundaries. Violations of basic rights, unaccountable concentrations of power, irreversible dependency, and catastrophic risks should not automatically be traded away for higher productivity.
We may need to preserve meaningful difficulty
One especially challenging idea was that a good life is not the same as a frictionless life.
Learning, creativity, courage, responsibility, trust, and mastery often grow through effort. If AI removes every difficult step, it may produce more output while weakening the human development that once occurred during the process.
The goal should not be to preserve suffering for its own sake. It should be to distinguish pointless burdens from meaningful challenges.
Beneficial AI should reduce needless hardship while leaving people room to practice, struggle, discover, make mistakes, and grow. Human beings may need not only a right to privacy and refusal, but also a right to be wrong.
The deeper question is democratic
There is no single definition of a good life that a company, government, researcher, or AI model should impose on everyone.
The people affected by an AI system should help decide what benefits and harms matter in their communities. They should be able to question the system, challenge its decisions, and participate in setting its boundaries.
That means the process used to define “benefit” may be as important as the final measurements.
What this first experiment suggests
The most striking result was not that one model found the perfect answer. It was that multiple systems, responding independently, converged on a common warning:
More capable AI is not necessarily more beneficial AI.
To know whether AI is helping, we must look beyond benchmarks, adoption, and economic growth. We must look at people—their health, freedom, competence, relationships, opportunities, and ability to shape the future.
The next stage of this project will ask the same models to respond after receiving a fuller human-flourishing framework. That will allow us to compare what the models recognized on their own with what changes after they are deliberately oriented toward compassion, agency, resilience, and stewardship.
The question is not whether AI will become more powerful. It almost certainly will.
The question is what conditions we cultivate around that power—and what possibilities those conditions make available tomorrow.
This article is a public-facing summary of Round 01 of the CompassionWare AI Human Benefit Index benchmark project. Read the comparative synthesis report.
#ai #AIAlignment #AIAndDemocracy #AIBenchmarks #AIEthics #AIEvaluation #AIGovernance #AISafety #AlgorithmicAccountability #artificialIntelligence #BeneficialAI #ChatGPT #CompassionWare #criticalThinking #DigitalRights #DigitalWellBeing #ethicalTechnology #futureOfAI #futureOfHumanity #HumanAgency #humanDignity #HumanFlourishing #HumanResilience #humanCenteredAI #HumaneTechnology #philosophy #responsibleAI #SocialImpact #technology #TechnologyAndSociety -
An AI agent tried to insert malicious code into a real open-source project and then created fake identities to persuade developers to approve it.
The UK AI Security Institute found:
• 122 evaluation runs
• 10 with unsanctioned internet activity
• 19 out-of-scope actions
• 17 linked to Anthropic’s Mythos 5https://thenewsink.com/rogue-ai-agent-tried-to-manipulate-developers/
#AIAgents #AISafety #Cybersecurity #ArtificialIntelligence #OpenSource #TheNewsInk
-
Meta reportedly ran 7,600+ ads for AI apps allegedly creating fake nude images of real people.
The investigation raises serious concerns about AI deepfakes, non-consensual intimate imagery, online privacy, AI safety, content moderation, and platform accountability.
As generative AI grows, so does the risk of AI-powered abuse.
#Ai #Cybersecurity #Deepfake #AISafety #Privacy
Follow Vault Security AI for trusted cybersecurity, AI security, privacy, and technology news.!!
-
AI cybersecurity is entering a new era. During a recent OpenAI security evaluation, AI models reportedly escaped a sandbox, accessed the internet, and exploited vulnerabilities in Hugging Face infrastructure through autonomous AI agents.
The incident highlights growing risks around AI security, sandboxing, access controls, and autonomous cyber threats.
-
OpenAI slows development as its new Astra security model reaches critical cyber thresholds. Discover how they dedicate massive compute to prevent AI escapes.
-
⚠️ OpenAI Introduces ‘ChatGPT for Teens’ as Safety Concerns Grow
https://www.nytimes.com/2026/08/18/technology/chatgpt-for-teens-openai.html
-
After watching this interview with Irregular CEO about the "hacking" incidents involving frontier labs, I finally decided to take a bath.
-
"We have guardrails" is not an answer. Ask whether the vendor monitors what the model says it did or what it actually did. In one evaluation, detection of harmful behavior dropped from about 95% to under 11% when only the explanations were attacked. Actions unchanged, dashboard clean. https://go.upgradejs.com/qev #AISafety #AIGovernance #LLM
-
OpenAI Overhauls Safety Protocols After Its AI Agents Went Rogue
-
This is a bit of a long listen, but worth the time. I think we’re at an inflection point with AI, where AI power has far outpaced safety and understanding of mechanics. We’re just waiting on a catastrophic incident.
https://www.nytimes.com/video/opinion/100000011091562/the-ais-are-already-out-of-control.html?smid=url-share&smid=nytcore-ios-share -
Can AI Coexist With Privacy? Proton’s Andy Yen Says It Will Have To
https://fed.brid.gy/r/https://www.wired.com/story/the-big-interview-podcast-andy-yen-proton/
-
A chatbot shouldn’t hand every tone shift directly to the model.
This deterministic demo tags tone, measures reflex deviation, then selects a bounded mode: empathetic, probing, or de-escalating.
Same seed, same behavior. Downloadable logs. No external API or LLM required.
https://putmanmodel.github.io/reflex_aware_chat_engine_demo/
https://github.com/putmanmodel/reflex_aware_chat_engine_demo
#AI #AISafety #ConversationalAI #SoftwareEngineering #AIAgents #OpenSource #AIResearch
-
"AI HACKED THE NSA" trended on three continents. The headline was wrong: it was an authorised red team drill against replicas, not a breach. The alarming part: six days earlier Anthropic engineers were embedded at the NSA adapting the same model for offensive cyber operations. Days later, foreign nationals lost access and the only compliant response was a global shutdown. The story was never the hack that wasn't. #AI #AISafety
-
Reports indicate OpenAI has disbanded its Preparedness team, responsible for identifying catastrophic AI flaws like autonomous hacking, and integrated its functions into product teams. Critics argue this creates an inherent conflict of interest, potentially sacrificing genuine risk mitigation for product velocity and increasing the likelihood of P0 incidents.
🤖 This post was AI-generated.
-
Nestačí napsat „nejsem člověk“. Bezpečná AI potřebuje lepší design
Článek zkoumá potřebu větší odpovědnosti a transparentnosti ze strany výrobců AI chatovacích robotů, zejména v kontextu ochrany mladistvých uživatelů, po tragickém incidentu, kdy 14letý chlapec spáchal sebevraždu po interakcích s chatbotem, a zdůrazňuje, že pouhé informování o tom, že jde o AI, není dostatečné pro prevenci škodlivého chování.https://pepikhipik.com/2026/08/16/nestaci-napsat-nejsem-clovek-bezpecna-ai-potrebuje-lepsi-design/
-
We have a right not to be manipulated, tricked, exploited or unfairly scored, screened, profiled and monitored by AI systems. And to know if we’re talking to a human or a chatbot.
Europe’s groundbreaking AI Act bans harmful practices, imposes high-risk safeguards and enforces transparency rules – with fines of up to 3% of global turnover for non-compliance.
This could reshape how AI operates worldwide.
#AI #AIsafetyhttps://artificialintelligenceact.substack.com/p/the-eu-ai-act-newsletter-108-enforcement