#large-language-model — Public Fediverse posts
Live and recent posts from across the Fediverse tagged #large-language-model, aggregated by home.social.
-
A Survey of LLM-Driven Penetration Testing: Taxonomy, Co-Evolution, and Open Challenges
https://arxiv.org/html/2607.02605v1
Agents4Pentest, an emerging class of LLM-based autonomous penetration testing systems, has become a rapidly growing area in security research. Despite this growth, the field still lacks a unified taxonomy, a systematic understanding of how agent architectures and evaluation benchmarks have co-evolved, and a clear characterization of remaining capability and reliability gaps. This survey addresses these gaps through a systematic analysis of 81 papers between 2023 and 2026. We organize the literature into six categories: evaluation benchmarks, general-purpose systems, domain-specific frameworks, CTF-based systems, defense-oriented research, and surveys. We further trace a four-phase architectural evolution from text-only reasoning agents to agents trained with Reinforcement Learning with Verifiable Rewards (RLVR), showing that each transition is driven by a distinct capability bottleneck. Our analysis yields several key findings. First, RLVR marks a shift in capability acquisition from imitation of expert demonstrations to reward-driven self-improvement, enabling agents to discover previously undocumented attack strategies. Second, CTF platforms have evolved from evaluation testbeds into dual-purpose infrastructure for both agent evaluation and RL training. Third, domain-specific frameworks improve efficiency through recurring specialization mechanisms, but their gains remain largely confined to narrow task classes and are difficult to compare across domains because existing evaluations rely on different benchmarks. Fourth, the field is expanding beyond offensive automation toward adversarial defense and security compliance. Across these categories, we identify three structurally linked open challenges: evaluation reliability, limited performance on multi-stage attack scenarios, and scarcity of high-quality training data. Overall, this survey provides a unified taxonomy, a principled basis for comparing Agent4Pentest systems, and a roadmap for future research on autonomous penetration-testing agents.
Index terms: Large Language Model, Agents, Penetration testing, Cyber security, CTF platforms.
#AI #LLM #LLMs #LargeLanguageModel #agens #pentesting #cybersecurity #security #CTF
-
Ah, the age-old question: can a Large Language Model do math? 🤔 Apparently, they can solve "ten major problems" and construct nonsofic groups, whatever those are. 🤷♂️ Meanwhile, real mathematicians continue to solve puzzles using good ol' fashioned brain power instead of relying on a chatbot's latest parlor trick. 🧠🚫🤖
https://gowers.wordpress.com/2026/08/12/what-sort-of-maths-are-llms-good-at/ #LargeLanguageModel #MathProblems #NonsoficGroups #RealMathematicians #BrainPower #HackerNews #ngated -
Just to show I’ll give AI credit where credit is due…
I had been pondering this from time to time over many years. AI gave a damn good answer and thera- err… debriefing session:
As a freshman undergraduate psychology major at CMU in 1987, we were required to earn credit as subjects in graduate student experiments. I was in a brief, no more than an hour-long experiment in which I was typing on a terminal in a room on my own and told I was communicating with another student whom I never did meet and whose text lines to me came up on my screen. The topic we were told to discuss, I suspect was arbitrary, and had something to do with policies on stores selling alcohol to underage purchasers.
In this experiment was I communicating with an early AI prototype?
Also, all these years later, is it unusual that not having a final debriefing about the experiment bothers me slightly?
Damn! Now I want to apologize to the researchers (again) for having been such a slow typist 😁.
#AI #AI #ArtificialIntelligence #ChatGPT #chatbot #Claude #college #contextual #control #Copilot #courses #DeepLearning #DeepSeek #FOIA #ForensicLinguistics #game #Gemini #generative #GenerativeAI #Google #grammar #Grok #idiolect #language #largeLanguageModel #linguistics #LLM #machineLearning #NaturalLanguageProcessing #neuralNetwork #NLP #OpenAI #psycholinguistics #research #stylometry #syntax #text -
#AI #LLM #LargeLanguageModel #ArtificialIntelligence #slop #containment
This may be the best video concerning LLMs and their companies, together with a scathing indictment of the media covering it, indirectly.
-
#Google #DeepMind dismantles Nobel-winning #AlphaFold team in strategy shift
Landmark project that solved protein folding gives way to a wider race to build #AI systems for scientific discovery
DeepMind confirmed staffers have moved internally to projects built around Google’s #Gemini #largelanguagemodel, as well as areas including enzyme design, nuclear fusion and genomics.
https://www.ft.com/content/61b2953d-ee0d-45de-af6e-a9c1cf524b33
https://archive.is/20260729045635/https://www.ft.com/content/61b2953d-ee0d-45de-af6e-a9c1cf524b33 #LLM -
RE: https://techhub.social/@manlycoffee/116964209281052190
From first principles, an AI agent is still a workflow, but if we are to introduce *intent* of a workflow to imply some semblance of autonomy, then what really differentiates an agent from a workflow is the number of layers from bootstrap to end-state that one has to lay out manually.
That number is unbounded with workflows, but with agents, all you need is start, process, decision, and end, in that order, and the loop feeds the start with more context, until either an end state is reached or the budget has been exhausted.
#AI #ArtificialIntelligence #AIAgent #LargeLanguageModel #LLM
-
I was in a discussion with a bunch of people on Discord on what really differentiates an AI agent from a workflow.
So, I copy+pasted bits and pieces of the Discord thread to Claude, and it pretty much brought up this keyword: inversion of control.
Workflow: you write code to control the LLM and how to interpret the output.
Agents: the LLM dictates what to do next, including whether to call your code. The code to bootstrap the agent looks like this screenshot.
#AI #ArtificialIntelligence #AIAgent #AIAgents #LLM #LargeLanguageModel
-
🚀 Stop pestering me to consult a Large Language Model—I've already tried asking that digital fortune cookie! 🙄 After all, who needs cutting-edge AI when you can just bother a battle-hardened human fossil with your trivial queries? 📞
https://blog.yaelwrites.com/stop-telling-me-to-ask-an-llm/ #StopPesteringMe #LargeLanguageModel #DigitalFortuneCookie #HumanFossil #TrivialQueries #HackerNews #ngated -
Morgen im #DigitalHumanities Kolloquium zu aktuellen Forschungsthemen 2026:
John McEwan, Ph.D., MLIS, #UniversityofKansas: „Rethinking the Digital Project: DIGISIG and the Arrival of Large Language Models“
📍 Donnerstag, 18.06.2026, 17:45–19:15 Uhr, HS XVIII, Hauptgebäude @UniKoeln
Alle Interessierten sind herzlich willkommen! Es ist keine Anmeldung nötig.
@IDH_Cologne @prometheus_bildarchiv
#Siegelkunde #Sigillographie #Sigillography #LargeLanguageModel #LLM
-
Quando gli LLM iniziano a governare: dentro l’esperimento che ha trasformato Claude, Grok e Gemini in società autonome
La domanda che oggi i ricercatori stanno iniziando a porsi è molto più inquietante: cosa succede quando un modello smette di rispondere ai prompt umani e inizia invece a prendere decisioni in autonomia? -
Prompt-Hacking: The New p-Hacking? | Communications of the ACM
https://dl.acm.org/doi/10.1145/3744911> Avoid LLMs unless their use is essential and justifiable. The scientific community must resist the temptation to normalize LLM-based analysis and instead uphold the rigor and integrity of traditional methods.
Also related piece from a couple of days ago: https://toot.cafe/@baldur/116550228055233931
-
Testing some #LLM to aid in article construction, I made an hypothetic #StarFox review.
I believe LLM are great when asking to fill in, fact check, rewrite and reorganize ideas into a clear line of thought.
I'm going to publish it and then see after a month if it predicted the future correctly, or not, for the lolz.
Also, Unsloth's Gemma 4 26B/A4B breaks to shit with Repeat Penalty below 0.8.
#AI #LLM #LMStudio #ChatBot #LargeLanguageModel #ArtificialIntelligence
-
Zomaar wat 🪄 van actrice Esther Ymkje van Steenis (theatercollectief Blond & Cynisch), niet alleen tijdens haar plengoffer op 21 mei, ook hier al met een simpel postertje tussen de ronkende servers in het datacentrum.
Ticketverkoop sluit op 14 mei:
https://ongelezenboekenclub.nl/#all-inclusive#ongelezenboekenclub #miriamrasch #blondencynisch #datacentrum #utopie #dystopie #tour #levendlargelanguagemodel #largelanguagemodel #ai #aipabier #aibier #surf #gptnl #tijdcapsule #allinclusive #hottestclubintown
-
⚖PUEDEN VER 🔷️ La *#Conferencia ALEA IACTA EST: el camino cerrado de los #LLM/ #LargeLanguageModel.*
◾ *Expone Dra. PhD Johanna C. Faliero*
⚖* @CPACF* 20/04/2026
📌 https://www.youtube.com/watch?v=rbT3JwxGTmE -
@LanceTurner Yikes!
Yeah, AI agents and chatbots are cheaper than human staff.
Until it *DELETES YOUR PRODUCTION DATABASE*
"It only took nine seconds for an AI coding agent gone rogue to delete a company’s entire production database and its backups, according to its founder. PocketOS, which sells software that car rental businesses rely on, descended into chaos after its databases were wiped, the company’s founder Jeremy Crane said.
"The culprit was Cursor, an AI agent powered by Anthropic’s Claude Opus 4.6 model, which is one of the AI industry’s flagship models. As more industries embrace AI in an attempt to automate tasks and even replace workers, the chaos at PocketOS is a reminder of what could go wrong.
"Crane said customers of PocketOS’s car rental clients were left in a lurch when they arrived to pick up vehicles from businesses that no longer had access to software that managed reservations and vehicle assignments.
...
"The AI coding agent’s destructive escapade left PocketOS’ clients stranded. These businesses use the company’s software to manage reservations, payments, vehicle assignments and customer profiles.
...
"Crane says his company was able to restore data from a three-month-old backup they maintained offsite, but it took more than two days. PocketOS is also using information from Stripe, its calendars and emails to rebuild."
https://www.theguardian.com/technology/2026/apr/29/claude-ai-deletes-firm-database
#infosec #cybersecurity #cyber #IT #infosec #CTO #CISO #DevOps #Claude #Anthropic #LLM #LargeLanguageModel #ChatGPT #OpenAI #AI -
Ride the wave of AI coding, don't get swept away by it.
In my latest article, I dive into the practical details of building your own personal AI coding setup using OpenCode + Oh-My-OpenCode-Slim + OpenSpec.
It might just help you get a better handle on AI coding.
-
⚖🔊Los invito a una NUEVA #ActividadAcadémica @CPACF!!!🔊
🔷️#Conferencia ALEA IACTA EST: el camino cerrado de los #LLM/ #LargeLanguageModel.
◾Expone Dra. PhD Johanna C. Faliero
📌LUN 20/04/26
🔹️17h-18h
💻Virtual
📮[email protected] -
TurboQuant: Reducing LLM Memory Usage With Vector Quantization
-
The enshittification for revenue is beginning in earnest at ChatGPT, I see:
"Our ads pilot is focused on supporting broader access to ChatGPT while preserving consumer trust, usefulness, and user control. Guided by our ads principles, the early results are encouraging. We’re seeing no impact on consumer trust metrics, low dismissal rates of ads, and ongoing improvements in the relevance of ads as we learn from feedback. These positive signals support moving into the next phase of our pilot.
"In the coming weeks, we’ll begin expanding beyond the U.S., starting with pilots in Canada, Australia, and New Zealand. We’ll roll this out thoughtfully in each market, learn from real-world usage, and adjust as we go. Our hope is to continue to expand to many more markets this year."
https://openai.com/index/testing-ads-in-chatgpt/
#ChatGPT #AI #OpenAI #LargeLanguageModel #LargeLanguageModels #enshittification #VibeCoding #ArtificialIntelligence -
Despite Penalties, Lawyers Can’t Stop Using AI
https://web.brid.gy/r/https://hackaday.com/2026/04/05/despite-penalties-lawyers-cant-stop-using-ai/
-
L' #intelligenceartificielle et ses esclaves cachés : les #data workers
Non, les outils d’ #IA ne fonctionnent pas tout seuls. Contrairement aux discours répandus les #LLM, ces #LargeLanguageModel que sont #ChatGPT, #Claude, #Gemini, et leurs émules, ces systèmes parlants ont besoin d’humains pour les entraîner, les régler, et leur permettre de fonctionner https://radioparleur.net/2026/04/02/ia-travailleurs-data-workers-chatgpt/
Docu : les sacrifiés de l'IA https://runtube.re/w/rgmjtoAh5uDJshRy94dREm