home.social

#airesearch — Public Fediverse posts

Live and recent posts from across the Fediverse tagged #airesearch, aggregated by home.social.

  1. So, more explanation of wp2shell recently just popped out.

    The vulnerability were found by GPT 5.6 Sol. By using modified prompt from how it found the solution of Cycle Double Cover conjecture.

    It was initially found a SQL Injection, but after asked again if it can be elevated to RCE, it confirms it in 4 hours.

    Technical explanation on the vulnearbility also can be found in this writeup, have a good read fellas.

    slcyber.io/research-center/exp

    #cybersecurity #infosec #security #wordpress #chatgpt #gptsol #wp2shell #airesearch #llm #vulnerability #vulnerabilityresearch

  2. Frontier models only is the latest expression of true AI belief, according to Gizmodo. The article explores how the tokenmaxxing approach to LLM optimisation has evolved and mutated into new forms. gizmodo.com/tokenmaxxing-didnt #AIagent #AI #GenAI #AIResearch

  3. Frontier models only is the latest expression of true AI belief, according to Gizmodo. The article explores how the tokenmaxxing approach to LLM optimisation has evolved and mutated into new forms. gizmodo.com/tokenmaxxing-didnt #AIagent #AI #GenAI #AIResearch

  4. Feyn AI has released SQRL, a text-to-SQL model family that inspects the database before writing queries. The flagship SQRL-35B-A3B achieves 70.6% execution accuracy on the BIRD benchmark, outperforming Claude Opus 4.6 at 68.77%. Available on Hugging Face. marktechpost.com/2026/07/19/fe #AIagent #AI #GenAI #AIResearch

  5. Feyn AI has released SQRL, a text-to-SQL model family that inspects the database before writing queries. The flagship SQRL-35B-A3B achieves 70.6% execution accuracy on the BIRD benchmark, outperforming Claude Opus 4.6 at 68.77%. Available on Hugging Face. marktechpost.com/2026/07/19/fe #AIagent #AI #GenAI #AIResearch

  6. Perplexity has released WANDR, an open benchmark with 500 evidence-heavy tasks for testing research agents. The benchmark evaluates whether agents can discover many qualifying entities and back each with cited evidence. Perplexity Search as Code leads at 0.363 soft F1. marktechpost.com/2026/07/19/pe #AIagent #AI #GenAI #AIResearch

  7. Perplexity has released WANDR, an open benchmark with 500 evidence-heavy tasks for testing research agents. The benchmark evaluates whether agents can discover many qualifying entities and back each with cited evidence. Perplexity Search as Code leads at 0.363 soft F1. marktechpost.com/2026/07/19/pe #AIagent #AI #GenAI #AIResearch

  8. The spinny 3D Ai Harness multi graph...
    ...why? Because we can.

    Engines that hit their compute limits are red. Some of the small wireframe engines have not had their "Power" evaluated yet, so the Harness router has them on the end of fall-through hierarchy.

    The skulll is a local guardrail free - model.

    You can clearly see the "consensus" artifact, thats the model connecting to 4 other models. All four must have the same anwser, otherwise a higher model steps in.

    Oh, and I bought some API time on togetherAi. Because its cheap and gives me access to GLS and Kimi models which are comparable with #Anthropic Opus

    Using those two models, I now can have a viable alternative to Opus/Sonnet with adequate tokens.

    #Vibecoding #Harness #AiResearch

  9. 🚀 Fastest-growing AI projects today

    1. One standout project "open-science," which has gained significant traction as an open-s...
    2. **ai4s-research/open-science**: Threpository a local-first, model-agnostic AI research...
    3. With its growth score of 53.47 and over 800 stars, it's clear that the project resonati...

    Full report → pullrepo.com/report/todays-ai-

    #AI #OpenSource #GitHub #Tech #AIResearch

  10. 🚀 Fastest-growing AI projects today

    1. One standout project "open-science," which has gained significant traction as an open-s...
    2. **ai4s-research/open-science**: Threpository a local-first, model-agnostic AI research...
    3. With its growth score of 53.47 and over 800 stars, it's clear that the project resonati...

    Full report → pullrepo.com/report/todays-ai-

    #AI #OpenSource #GitHub #Tech #AIResearch

  11. Three Chinese AI labs have released open-weight MoE models. Kimi K3 leads on benchmarks but costs more to serve. DeepSeek V4 Pro is cheapest at 0.04 USD per task. GLM-5.2 balances performance and cost. marktechpost.com/2026/07/18/ki #AIagent #AI #GenAI #AIResearch

  12. Three Chinese AI labs have released open-weight MoE models. Kimi K3 leads on benchmarks but costs more to serve. DeepSeek V4 Pro is cheapest at 0.04 USD per task. GLM-5.2 balances performance and cost. marktechpost.com/2026/07/18/ki #AIagent #AI #GenAI #AIResearch

  13. RT @dunik_7: TRANSLASATION: Ein Labor der Tsinghua-Universität hat ein Projekt auf GitHub veröffentlicht, das einen H100-Rack im Wert von 400.000 US-Dollar durch eine einzelne 24-GB-Grafikkarte ersetzt. Das Projekt heißt ktransformers, und der Trick ist fast schon lächerlich einfach: Die Experten, die Sie tatsächlich nutzen, bleiben auf der GPU, während die anderen auf der CPU warten, bis sie benötigt werden. / DeepSeek-V3 und R1 mit 139K Kontext in 24GB VRAM / bis zu 28-fache Geschwindigkeitssteigerung gegenüber dem Standard-Setup / Fine-Tuning von DeepSeek-V3 über vier RTX 4090 statt eines Rechenzentrums / entwickelt vom MADSys-Labor der Tsinghua-Universität, nicht von einem Startup mit einer Landing Page. Apache 2.0, bereits über 17.000 Sterne. - github.com/kvcache-ai/ktransfo merken.

    mehr auf Arint.info

    #AIResearch #DeepSeekV3 #ktransformers #MachineLearning #OpenSource #TsinghuaUniversity #arint_info

    https://x.com/dunik_7/status/2078065378563887290#m

  14. Zyphra has released ZUNA1.1, an open-source EEG foundation model that processes brain signals from 0.5 to 30 seconds. The 380M parameter model reconstructs and denoises EEG data across arbitrary electrode layouts, advancing brain-computer interface research. marktechpost.com/2026/07/17/zy #AIagent #AI #GenAI #AIResearch

  15. Zyphra has released ZUNA1.1, an open-source EEG foundation model that processes brain signals from 0.5 to 30 seconds. The 380M parameter model reconstructs and denoises EEG data across arbitrary electrode layouts, advancing brain-computer interface research. marktechpost.com/2026/07/17/zy #AIagent #AI #GenAI #AIResearch

  16. NVIDIA has unveiled Nemotron 3 Embed, an open-source embedding model collection. The 8 billion parameter version tops the RTEB benchmark for retrieval tasks, marking a significant advancement in AI infrastructure for enterprise search and RAG applications. marktechpost.com/2026/07/17/nv #AIagent #AI #GenAI #AIResearch

  17. NVIDIA has unveiled Nemotron 3 Embed, an open-source embedding model collection. The 8 billion parameter version tops the RTEB benchmark for retrieval tasks, marking a significant advancement in AI infrastructure for enterprise search and RAG applications. marktechpost.com/2026/07/17/nv #AIagent #AI #GenAI #AIResearch

  18. 🎥 Missed the 8th Weizenbaum Conference?

    Both keynote lectures are now available on our YouTube channel.

    🤖 Under the theme "Generative AI and Society: What is at stake?", more than 300 researchers and experts discussed the societal implications of generative AI.

    Watch:

    ▶️ Nick Srnicek – "The Rise of Silicon Empires"
    lnkd.in/d3kN6-M3

    ▶️ Alexander Campolo – "Zero Shot World: On the Political Logics of Generative AI"
    lnkd.in/dMy9AZk9

    #GenerativeAI #DigitalSociety #AIResearch

  19. 🎥 Missed the 8th Weizenbaum Conference?

    Both keynote lectures are now available on our YouTube channel.

    🤖 Under the theme "Generative AI and Society: What is at stake?", more than 300 researchers and experts discussed the societal implications of generative AI.

    Watch:

    ▶️ Nick Srnicek – "The Rise of Silicon Empires"
    lnkd.in/d3kN6-M3

    ▶️ Alexander Campolo – "Zero Shot World: On the Political Logics of Generative AI"
    lnkd.in/dMy9AZk9

    #GenerativeAI #DigitalSociety #AIResearch

  20. China has a new top model. Moonshot AI's Kimi K3 is very good — but the hype may be getting ahead of reality. The 2.8-trillion-parameter open model is claimed to beat Claude Fable 5 and GPT 5.6 in some benchmarks. platformer.news/kimi-k3-launch #AIagent #AI #GenAI #AIResearch

  21. China has a new top model. Moonshot AI's Kimi K3 is very good — but the hype may be getting ahead of reality. The 2.8-trillion-parameter open model is claimed to beat Claude Fable 5 and GPT 5.6 in some benchmarks. platformer.news/kimi-k3-launch #AIagent #AI #GenAI #AIResearch

  22. Moonshot AI has released Kimi K3, a 2.8 trillion parameter open Mixture-of-Experts model with a 1 million token context window. The model uses Kimi Delta Attention and is the first open 3T-class model. marktechpost.com/2026/07/16/mo #AIagent #AI #GenAI #AIResearch

  23. Moonshot AI has released Kimi K3, a 2.8 trillion parameter open Mixture-of-Experts model with a 1 million token context window. The model uses Kimi Delta Attention and is the first open 3T-class model. marktechpost.com/2026/07/16/mo #AIagent #AI #GenAI #AIResearch

  24. 🚨 Breaking News: Yet another attempt to identify LLM-generated texts with classical ML techniques! 🤯 Spoiler: it’s like using a magnifying glass to find a needle in a haystack. 🔍⚠️ Enjoy the #TLDR, because who needs 10 useless subheadings just to say "we're still guessing"? 😂
    blog.lyc8503.net/en/post/llm-c #BreakingNews #LLM #TextAnalysis #MLtechniques #AIresearch #HackerNews #ngated

  25. 🚨 Breaking News: Yet another attempt to identify LLM-generated texts with classical ML techniques! 🤯 Spoiler: it’s like using a magnifying glass to find a needle in a haystack. 🔍⚠️ Enjoy the #TLDR, because who needs 10 useless subheadings just to say "we're still guessing"? 😂
    blog.lyc8503.net/en/post/llm-c #BreakingNews #LLM #TextAnalysis #MLtechniques #AIresearch #HackerNews #ngated

  26. 🚨 Breaking News: Yet another attempt to identify LLM-generated texts with classical ML techniques! 🤯 Spoiler: it’s like using a magnifying glass to find a needle in a haystack. 🔍⚠️ Enjoy the #TLDR, because who needs 10 useless subheadings just to say "we're still guessing"? 😂
    blog.lyc8503.net/en/post/llm-c #BreakingNews #LLM #TextAnalysis #MLtechniques #AIresearch #HackerNews #ngated

  27. 🚨 Breaking News: Yet another attempt to identify LLM-generated texts with classical ML techniques! 🤯 Spoiler: it’s like using a magnifying glass to find a needle in a haystack. 🔍⚠️ Enjoy the #TLDR, because who needs 10 useless subheadings just to say "we're still guessing"? 😂
    blog.lyc8503.net/en/post/llm-c #BreakingNews #LLM #TextAnalysis #MLtechniques #AIresearch #HackerNews #ngated

  28. 🚨 Breaking News: Yet another attempt to identify LLM-generated texts with classical ML techniques! 🤯 Spoiler: it’s like using a magnifying glass to find a needle in a haystack. 🔍⚠️ Enjoy the #TLDR, because who needs 10 useless subheadings just to say "we're still guessing"? 😂
    blog.lyc8503.net/en/post/llm-c #BreakingNews #LLM #TextAnalysis #MLtechniques #AIresearch #HackerNews #ngated