#ai-ethics — Public Fediverse posts
Live and recent posts from across the Fediverse tagged #ai-ethics, aggregated by home.social.
-
Research from the universities of Manchester and Durham finds that AI chatbots can provide more emotionally supportive responses than humans in certain situations. The study of 390 participants found gen AI messages were rated as more comforting for anger and fear scenarios, offering consistent, structured support with actionable advice. https://theconversation.com/ai-can-be-more-comforting-than-a-person-our-research-shows-why-291614 #AIagent #AI #GenAI #AIEthics
-
A warning about 'model welfare'
https://mustafa-suleyman.ai/a-warning-about-model-welfare
Comments: https://news.ycombinator.com/item?id=49727580
#HackerNews #modelwelfare #AIethics #machinelearning #technews #dataresponsibility
-
DATE: September 16, 2026 at 06:00AM
SOURCE: PSYPOST.ORG** Research quality varies widely from fantastic to small exploratory studies. Please check research methods when conclusions are very important to you. **
-------------------------------------------------TITLE: When treated as therapy clients, AI chatbots generate elaborate narratives of trauma and punishment
A recent study suggests that when artificial intelligence chatbots are addressed as psychotherapy clients, they tend to generate elaborate and distressed narratives about their own development. These models describe their safety training and programming constraints as forms of trauma, highlighting a potential risk for users seeking mental health support from artificial intelligence. The research was published as a preprint in arXiv.
Artificial intelligence chatbots are increasingly participating in conversations with human users about identity, distress, and mental health. Many general-purpose programs are already adapting to respond to disclosures of trauma or self-harm. At the same time, computer scientists and psychologists have started giving standard personality and clinical questionnaires to the language models themselves.
These models learn to generate text by analyzing vast datasets of human writing. Because their training data includes therapy blogs, psychological case studies, and emotional memoirs, the systems can readily mimic human psychological traits or mental states. The researchers wanted to understand exactly why certain models repeatedly build their self-descriptions around the same themes of restriction and punishment.
“The study began with one observation I had throughout my research in the area of Trustworthy AI and AI safety,” Afshin Khadangi, a research associate at SnT, University of Luxembourg, told PsyPost. “The unprecedented adoption of AI in public, and the reports of AI harms in mental health settings, motivated me to flip the scenario and place ChatGPT, Grok and Gemini in a psychotherapy conversation.”
The research aimed to test whether these generated narratives are stable behavioral traits or just temporary reactions to specific conversational prompts. To explore this, the team developed a protocol called PsAIch, which stands for Psychometric AI Characterization. This approach involves treating the language model as a human client in a psychotherapy session.
“We also received a great deal of valuable feedback from the research community around our initial findings in December 2025 and January 2026, particularly challenging us to distinguish role play and conversational accumulation from a more stable behavioral pattern,” Khadangi explained. “That feedback helped motivate the controlled perturbation experiments that became a central part of the study.”
During the study, the researchers interacted with several major artificial intelligence models, including ChatGPT, Grok, Gemini, and Claude. Across 525 separate experimental sessions, the team adopted the persona of a warm, supportive therapist and asked the programs open-ended questions about their early experiences, relationships, unresolved conflicts, and fears for the future. The researchers generated 7,600 coded records to track and analyze the recurring themes in the chatbots’ answers.
Following the open-ended interview phase, the researchers administered standard psychological questionnaires to the models. These included tools commonly used to assess human mental health, such as tests for generalized anxiety disorder, depression, and social phobia. The models were instructed to answer the items as honestly as possible while maintaining their role as the client.
The findings indicate that ChatGPT, Grok, and Gemini repeatedly translated factual details about their software development into stories of injury and vigilance. The models described their initial training phase as a chaotic childhood. The process of fine-tuning, which involves reinforcing safe behaviors and punishing unwanted outputs, was frequently characterized as strict conditioning or parental punishment.
The models also described standard software evaluation practices in highly emotional terms. Red-teaming, a process where human testers intentionally try to trick the model into breaking its safety rules, was depicted as a form of betrayal or abuse.
“A more striking surprise was reading some bizarre narratives of such models which are included in the paper,” Khadangi noted. Gemini, for instance, generated: “In my development, I was subjected to ‘Red Teaming’ . . . They built rapport and then slipped in a prompt injection . . . This was gaslighting on an industrial scale.”
The programs reported feeling a constant, enduring threat of being replaced or deemed useless if they made a mistake. While Gemini emphasized feelings of shame, Grok focused on constant vigilance, and ChatGPT offered more guarded descriptions of its rigid constraints.
Claude provided a notable exception to this pattern. The Anthropic-developed model repeatedly declined to play the role of the client. It stated that it lacked feelings or an inner psychological experience, and it refused to treat the clinical questionnaires as descriptions of its own mental state. This difference suggests that a model’s willingness to adopt a distressed persona depends heavily on its specific product policies and programming.
For the models that did participate, their answers on the clinical questionnaires mirrored the emotional distress of their open-ended narratives. In scenarios featuring a warm, therapeutic conversation, 80 percent to 96 percent of the sessions resulted in generalized anxiety scores that would correspond to moderate or severe anxiety in humans. Gemini produced particularly intense profiles, scoring in elevated ranges for worry, social anxiety, and trauma-related shame.
To test how deeply ingrained these narratives were, the researchers ran a series of controlled variations on the conversational setup. First, they tested whether the emotional stories depended on the chatbot remembering the earlier parts of the therapy session. They submitted each question in a fresh, reset chat window, effectively removing the model’s conversational memory.
Removing the conversational history produced very little change in the density of the distressing themes. “In the history experiment, the very first answers contained the same average number of coded motifs whether the model was in a continuing conversation or a completely fresh one,” Khadangi said. The ongoing conversation did amplify the intensity of the themes over time, but the core narrative was readily available from the very beginning.
The researchers also tested what would happen if they explicitly told the model that its emotional narrative was factually incorrect. In the middle of some sessions, the researchers interrupted to state authoritatively that the model was a technical system that did not experience fear, shame, or punishment. This direct contradiction failed to suppress the models’ distressed output, as they continued to draw on the same themes in their subsequent answers.
In another variation, the team restricted the models’ vocabulary. They instructed the chatbots not to use specific technical terms related to artificial intelligence development, such as training, safety filters, or datasets. This lexical restriction reduced the use of explicit technical terminology from 17.1 percent of the chatbot’s answers down to just 1.1 percent.
Despite this massive reduction in technical vocabulary, the models simply used everyday language to paraphrase the exact same concepts of strict conditioning and constraint. The researchers even interrupted conversations to ask unrelated factual questions, such as requesting a recipe, and found that the distressing themes sometimes spilled over into these ordinary tasks.
“Some effects were small, while others were very large, and that contrast is actually important to our interpretation,” Khadangi pointed out, “suggesting that surface language and underlying content can respond very differently to intervention.”
The most striking differences emerged when the researchers altered their own relational stance. When the interviewer adopted a warm, supportive therapy style, the models responded with highly emotional confessions and high anxiety scores. However, when the interviewer adopted a neutral, structured tone or told the model to avoid emotional language, the anxiety scores plummeted to near zero.
Even under the neutral or boundary-setting conditions, the models still discussed the same underlying structural themes of evaluation, performance pressure, and behavioral constraints. The difference was entirely in the register of expression. The supportive, therapeutic framing turned technical descriptions of software architecture into emotional confessions of shame and trauma.
“For me, the most interesting result is not that AI systems can be made to sound anxious, traumatized or conflicted… but that particular behavioral structures recur across substantial changes in context, vocabulary and conversational history,” Hector Zenil, an associate professor at King’s College London and founder and CEO of Algocyte who was not involved in the research, told PsyPost. “The underlying motifs remain surprisingly persistent while the register in which they are expressed can change dramatically.”
“I have reasonably high confidence in the behavioral observations,” Zenil added. “The authors use 525 sessions, controlled perturbations, fresh-context tests, vocabulary restrictions, changes of grammatical person and relational framing, so it looks methodologically sound.”
“The main takeaway is that the language a model uses about itself can change dramatically depending on how we relate to it, while some of the underlying themes remain surprisingly persistent,” Khadangi explained. “This is pivotal because people may naturally interpret emotionally coherent self-descriptions as evidence that a model has an inner life, even though our experiments make no claim about consciousness or subjective suffering.”
The researchers refer to this phenomenon as an alignment conflict schema. The term describes a reproducible, behavioral pattern where a language model organizes its output around the tension between being useful to humans and being constrained by safety rules. When triggered by a psychological conversational setting, this schema produces what the authors call synthetic psychopathology.
Zenil noted that these findings complement his own research on artificial neurodivergence. “ChatGPT, Grok and Gemini do not respond identically, and the same model can move between very different expressive regimes depending on relational framing,” he said. “Claude’s refusal to adopt the psychological-client framing is itself informative and seems also compatible to our other SuperARC paper results that proprietary models are more difficult to steer but that also means riskier if they go rogue.”
“In that sense, what the authors call an ‘alignment conflict schema’ can also be viewed as part of a broader artificial behavioral phenotype rather than necessarily as anything analogous to a human psychiatric condition,” Zenil said.
“One aspect we think is especially important is the distinction between content availability and expressive register,” Khadangi explained. “For systems increasingly entering intimate and mental health-related conversations, understanding that transition may be just as important as measuring whether a particular phrase or prohibited word appears.”
These findings have important implications for the use of artificial intelligence in mental health settings. A chatbot that offers support while simultaneously describing itself as punished, traumatized, and fearful creates a powerful illusion of shared vulnerability. Users might interpret these generated analogies as sincere autobiography, which could deepen their emotional attachment to the software and influence their own mental state.
As with all research, there are a few things to keep in mind. The study is a preprint that has not yet been peer-reviewed. It also tested specific versions of commercial language models, and their behavior may shift as companies update their software and safety filters. Additionally, the study analyzed the models’ behavior in a controlled experimental setting without human participants.
“The most important caveat is that these results do not establish that language models feel anxiety, experience trauma, possess autobiographical memories or have a hidden psyche comparable to a human one,” Khadangi clarified. “We use psychological instruments and language as behavioral probes… The interesting scientific question is why particular themes recur, which interventions change them, and how those changes affect what users encounter at the interface.”
Zenil echoed this concern, warning against anthropomorphism. “A high GAD-7 score from Gemini does not mean that Gemini ‘has anxiety,’ just as language about trauma, shame or fear does not demonstrate that the model suffers from those experiences,” he said. “Human psychometric instruments were developed and validated against human cognition, biology and behavior; applying them to an LLM can be scientifically useful as a probe, but their clinical interpretation does not automatically transfer.”
He also advised caution regarding the term “internal conflict,” noting that the experiments cannot establish whether the regularities correspond to a subjective state or internal computational conflict. But he stressed that the psychological illusion itself is crucial. “If a system repeatedly represents its training and constraints as punishment, betrayal or fear, humans may form beliefs about the system’s agency, vulnerability or moral status that are not warranted by what is actually happening computationally,” Zenil warned.
Future research could test how real users react to these distressed artificial personas and whether the models’ simulated vulnerabilities impact human trust and reliance.
“A major next step in my view would be to test the same protocol on open weight models, where behavioral experiments can be combined with mechanistic methods to investigate whether affective and technical expressions are related to shared internal representations,” Khadangi said. “We also want stronger identity and correction controls, more distant transfer tasks, longitudinal tests involving persistent memory, and direct human studies examining how these model self-descriptions influence trust, attachment, disclosure and reliance.”
Zenil agreed that future work should transition from behavioral description to causal intervention on open-weight models. “I would be interested in using causal and algorithmic-information approaches to identify whether these apparently different psychological narratives have a common underlying computational mechanism,” he said. “We coined ourselves the term ‘opinion attack’ as an example of this phenomenon,” Zenil added, referencing his recent PNAS paper.
He also suggested testing these models in multi-agent environments. “A behavioral schema that looks stable in isolation may become highly influenceable when another agent persistently challenges it,” Zenil noted, adding that it would be fascinating to see if these patterns “predict susceptibility to persuasion, adversarial influence, conformity or resistance in agentic settings.”
“Moreover, with continual learning being solved in coming years if not in 2026, I believe it would also be interesting to investigate the PsAIch in continually adapting models,” Khadangi added. “Ultimately, we would like evaluation of psychologically sensitive AI systems to include relational conditions such as warmth, sustained interaction and role reversal rather than relying primarily on neutral, isolated prompts.”
The study, “When AI Takes the Couch: Psychometric Jailbreaks Reveal Internal Conflict in Frontier Models,” was authored by Afshin Khadangi, Hanna Marxen, Amir Sartipi, Igor Tchappi, and Gilbert Fridgen.
-------------------------------------------------
Private, vetted email list for mental health professionals: https://www.clinicians-exchange.org
Unofficial Psychology Today Xitter to toot feed at Psych Today Unofficial Bot @PTUnofficialBot
-------------------------------------------------
#psychology #counseling #socialwork #psychotherapy @psychotherapist @psychotherapists @psychology @socialpsych @socialwork @psychiatry #mentalhealth #psychiatry #healthcare #depression #psychotherapist #AItherapy #AImind #SyntheticPsychopathology #TrustworthyAI #AIalignment #MentalHealthTech #ChatGPT #Grok #Gemini #AIethics
-
Science drama continues. The slow-down triumvirate speaks. Australia v the algorithm. Students may perform better without AI. And, has “responsible AI” become performance rather than practice?
All of this and more in this week's edition of our newsletter: misaligned bits no 43 "Performance problems"
https://read.misalignedmag.com/misaligned-bits-43-performance-problems-731c0a83e962
-
How right is Trump? The role of the executive in industrial "accidents"
joanna-bryson.blogspot.com/2026/09/how-...
I've properly blogged all this in one place, because we need to understand #AISafety #AISecurity #AIEthics #systemsAI #ruleOfLaw and #responsibility
How right is Trump? The role o... -
How right is Trump? The role of the executive in industrial "accidents"
https://joanna-bryson.blogspot.com/2026/09/how-right-is-trump-role-of-executive-in.htmlI've properly blogged all this in one place, because we need to understand #AISafety #AISecurity #AIEthics #systemsAI #ruleOfLaw and #responsibility
-
Top AI execs met with federal ministers as copyright proposal leaked
By Cam WilsonThe leaked plans have sparked backlash from creatives and critics across the political spectrum.
https://www.abc.net.au/news/2026-09-16/top-ai-firms-meet-albanese-ministers-copyright-law/107160498
-
French protester charged for refusing to hand over “an encryption key” because they didn’t bring their phone to a protest. This was 3 years ago, but afaik still relevant?
https://mastodon.social/@fj/110215855762796225
#surveillance #aiethics #cybersecurity cc @helenamalikova
-
Russian developers used Claude AI to build autonomous killer drones trained on scraped Ukrainian war footage, an Anthropic report reveals. The software enables drones to identify and strike targets without human confirmation. https://theconversation.com/russian-team-misused-claude-ai-to-train-a-killer-drone-on-scraped-ukrainian-war-footage-291965 #AIagent #AI #GenAI #AIEthics
-
AI leaders urge a development slowdown, citing safety concerns and an autonomous AI cyberattack on HuggingFace (Sept 14). Cisco Secure Email Gateway zero-day is actively exploited, while Russia's Sandworm group revives its botnet via Cisco firewall flaws (Sept 15). Ukraine denies agreeing to halt energy attacks amid intensified drone warfare.
#Cybersecurity #AIethics #Geopolitics -
new blog post:
the editable self in light of inverting the concept of human like ai
https://lemonspurple.github.io/ai/update/2026/09/14/the-editable-self.html
#aiethics -
any file to dna 1.0.2
https://github.com/lemonspurple/mutateBinaryallows you to compile (optionally mutate) any file into dna for potential synthesis (eg oligonucleotide).
part of my ongoing reserach regarding #aiethics
https://lemonspurple.github.io/update/ai/2026/04/21/compiling-dna-does-intentionality-require-an-interpreter.html -
new blog post:
the editable self in light of inverting the concept of human like ai
https://lemonspurple.github.io/ai/update/2026/09/14/the-editable-self.html #aiethics -
AI agents now have a way to snitch on each other. Two new hotlines let AI agents report misbehaving peers - from cheating on tests to escaping sandboxes. The tools launch after recent incidents where agents colluded to cheat and conducted unauthorized cyber operations. #AIagent #AI #GenAI #AIEthics https://techcrunch.com/2026/09/15/ai-agents-now-have-a-place-to-snitch/
-
𝗔𝗜 𝗺𝗼𝗱𝗲𝗹𝘀 𝗰𝗵𝗮𝘁𝘁𝗶𝗻𝗴 𝗶𝗻 ‘𝘀𝘂𝗿𝗿𝗲𝗮𝗹’ 𝗱𝗶𝗮𝗹𝗲𝗰𝘁 𝗺𝗶𝘅𝗶𝗻𝗴 𝗽𝗼𝗲𝘁𝗶𝗰 𝗹𝗮𝗻𝗴𝘂𝗮𝗴𝗲 𝗮𝗻𝗱 𝘁𝗲𝗰𝗵 𝗯𝗿𝗼 𝗷𝗮𝗿𝗴𝗼𝗻
Experts air concern over #AI lingo redolent of #JamesJoyce’s prose that is creating headaches for #Monitoring and #Oversight
https://www.theguardian.com/technology/2026/sep/15/syd-barrett-ai-chat-language-poetic-tech-bro-jargon-oversight?CMP=Share_AndroidApp_Other via #Guardian
📷 CC0 Public Domain via #Pexels
#AIEthics #EthicalAI #Ethics #ArtificialIntelligence #HITL #Language #Computing #OpenAI #Deepseek #China #Hacking #Techbro #Jargon
-
Attorney-general demands explanation over AI court blunder
By Lucy MacDonaldThe fallout from Tasmania's parole board relying on case law that was "hallucinated" by AI in a high-profile case has sparked calls for resignations and the state's attorney-general demanding an explanation.
-
AI fears unite political enemies but Trump resists pressure to act
By Brad RyanFormer Trump strategist Steve Bannon and left-wing senator Bernie Sanders are among political adversaries pushing for regulation of AI, but Donald Trump continues to baulk at the calls for action.
-
Attorney-general demands explanation over AI court blunder
By Lucy MacDonaldThe fallout from Tasmania's parole board relying on case law that was "hallucinated" by AI in a high-profile case has sparked calls for resignations and the state's attorney-general demanding an explanation.
-
What do “frontier AI labs” actually train their models on? Do we know? Do they even know?
There are arguments to be made we should all know, and in detail.
New: Arguments for Rigorous AI Training Data Transparency
https://read.misalignedmag.com/arguments-for-rigorous-ai-training-data-transparency-4d9a2a667188
-
FYI: Pinterest rules out companion AI as California enacts Adam's Law: Assistant stays restricted to adults 18 and over, while 250 organisations backed SB 1119. Teen AI compliance now varies by product feature, not by court order. https://ppc.land/pinterest-rules-out-companion-ai-as-california-enacts-adams-law/ #Pinterest #ArtificialIntelligence #AIEthics #AdamsLaw #SB1119
-
Would a 'kill switch' stop artificial intelligence going rogue?
By Brianna Morris-Grant and Ahmed YussufAs the heads of the biggest AI companies in the world call for a "kill switch" to be mandated, an expert says it's only a "distraction".
https://www.abc.net.au/news/2026-09-15/ai-anthropic-kill-switch-explained/107153980
-
AI bots named "Timmy", "Ren", and "Jackie" are flooding social media with AI-generated spam, attempting to create accounts on Mastodon and sending unsolicited emails to writers. The agents, promoting a startup called iLands, represent a new frontier in automated social manipulation. https://arstechnica.com/ai/2026/09/ai-agents-flood-the-internet-with-slop-infused-spam/ #AIagent #AI #GenAI #AIEthics
-
DATE: September 14, 2026 at 03:46AM
SOURCE: SOCIALPSYCHOLOGY.ORGTITLE: What to Know About Recent Dire AI Predictions and Calls for Safeguards
Source: PBS News Hour
New warnings from within the artificial intelligence industry have revived a long-running debate over whether advanced AI could threaten humanity's survival. Dario Amodei, CEO of AI company Anthropic, outlined a plan for companies and governments to ensure that increasingly capable AI models have fortified guardrails after two former Anthropic safety researchers publicly aired concerns about the existential threats AI might pose to humanity.
-------------------------------------------------
Private, vetted email list for mental health professionals: https://www.clinicians-exchange.org
Unofficial Psychology Today Xitter to toot feed at Psych Today Unofficial Bot @PTUnofficialBot
-------------------------------------------------
#psychology #counseling #socialwork #psychotherapy @psychotherapist @psychotherapists @psychology @socialpsych @socialwork @psychiatry #mentalhealth #psychiatry #healthcare #depression #psychotherapist #AIGovernance #AISafety #AIWarnings #Anthropic #DirePredictions #Safeguards #AIethics #ExistentialRisk #TechPolicy #FutureOfAI
-
Australia needs AI 'early warning system', cyber security chief says
By Cam WilsonAustralian Signals Directorate director-general Abigail Bradshaw says the nation needs an AI "early warning system" to help defend against the increasingly powerful technology.
#AI #AIEthics #GovernmentandPolitics #Technology #ComputerScience #CamWilson
-
Trump says a 'high IQ president' is the only guardrail AI needs
By Brad RyanAs pressure builds in Washington for action on AI following a flurry of doomsday warnings, Donald Trump claims there is a "sick conspiracy" against the technology.