home.social

#gpt52 — Public Fediverse posts

Live and recent posts from across the Fediverse tagged #gpt52, aggregated by home.social.

fetched live
  1. Как читать новости об ИИ и отличать прорыв от пресс-релиза. И как относиться к заголовкам про «ИИ отнимет работу»

    Новости об ИИ выходят быстрее, чем успеваешь их переварить: релизы моделей, таблицы бенчмарков, заявления про "революцию" и "конец профессий". Эта статья научит,что проверять, когда выходит новая модель, как читать бенчмарки, на что смотреть в model/system card , чтобы понимать реальный смысл анонса, чем open-weight отличается от закрытых моделей и почему это влияет на рынок. А заодно, как читать без паники и самообмана статьи вроде "ИИ отнимет у вас работу".

    habr.com/ru/articles/1003130/

    #нейросети #искусственный_интеллект #LLM #бенчмарки #Claude_Sonnet_46 #Gemini_31_Pro #GPT52 #SWEbench #ARCAGI2 #сравнение_моделей_ИИ

  2. GPT-5.2가 이론물리학 난제 풀었다, 15년 미해결 글루온 상호작용 발견

    GPT-5.2가 15년간 미해결로 남았던 글루온 입자 상호작용 문제를 해결하고 수학적 증명까지 완성했습니다. AI 보조 과학 연구의 새로운 전환점을 소개합니다.

    aisparkup.com/posts/9277

  3. GPT-5.2가 이론물리학 난제 풀었다, 15년 미해결 글루온 상호작용 발견

    GPT-5.2가 15년간 미해결로 남았던 글루온 입자 상호작용 문제를 해결하고 수학적 증명까지 완성했습니다. AI 보조 과학 연구의 새로운 전환점을 소개합니다.

    aisparkup.com/posts/9277

  4. GPT-5.2 proved gluon interactions physicists presumed to vanish actually exist in special conditions. OpenAI AI spent 12 hours deriving the formula verified by Harvard, Cambridge experts. AdwaitX explains the breakthrough 🔗 #AdwaitX #GPT52 #AIScience #Physics

    adwaitx.com/gpt-5-2-theoretica

  5. GPT-5.2 proved gluon interactions physicists presumed to vanish actually exist in special conditions. OpenAI AI spent 12 hours deriving the formula verified by Harvard, Cambridge experts. AdwaitX explains the breakthrough 🔗 #AdwaitX #GPT52 #AIScience #Physics

    adwaitx.com/gpt-5-2-theoretica

  6. OpenAI hat Deep Research auf GPT-5.2 migriert. Das Update bringt Site-Specific Search, um Quellen auf verifizierte Domains zu beschränken. Zudem erlauben App Connectors den Zugriff auf Daten aus Drittanbieter-Anwendungen. Nutzer können laufende Suchprozesse nun in Echtzeit stoppen oder korrigieren. Die Ausgabe erfolgt in einer neuen Vollbildansicht. #OpenAI #DeepResearch #GPT52
    all-ai.de/news/news26top/opena

  7. OpenAI hat Deep Research auf GPT-5.2 migriert. Das Update bringt Site-Specific Search, um Quellen auf verifizierte Domains zu beschränken. Zudem erlauben App Connectors den Zugriff auf Daten aus Drittanbieter-Anwendungen. Nutzer können laufende Suchprozesse nun in Echtzeit stoppen oder korrigieren. Die Ausgabe erfolgt in einer neuen Vollbildansicht. #OpenAI #DeepResearch #GPT52
    all-ai.de/news/news26top/opena

  8. Snowflake와 OpenAI 2억 달러 파트너십, 엔터프라이즈 AI 경쟁의 새로운 판도

    Snowflake와 OpenAI의 2억 달러 파트너십이 엔터프라이즈 AI 시장을 어떻게 재편하는지 분석합니다. 데이터를 옮기지 않고 AI를 쓰는 네이티브 통합의 의미를 소개합니다.

    aisparkup.com/posts/9014

  9. Snowflake와 OpenAI 2억 달러 파트너십, 엔터프라이즈 AI 경쟁의 새로운 판도

    Snowflake와 OpenAI의 2억 달러 파트너십이 엔터프라이즈 AI 시장을 어떻게 재편하는지 분석합니다. 데이터를 옮기지 않고 AI를 쓰는 네이티브 통합의 의미를 소개합니다.

    aisparkup.com/posts/9014

  10. ИИ против кандидата: как пройти собеседование, если HR — бот

    Ваше следующее собеседование начнется не в Zoom, а в интерфейсе, напоминающем чат с техподдержкой. Ваш HR-менеджер не будет знать, что у вас сегодня болит голова или что вчера вы сдали крутой проект. Он этого не знает, потому что он - это оно. Алгоритм, обученный на миллионах резюме и диалогов, цель которого - не понять вас, а отфильтровать. Добро пожаловать в 2026 год. Эпоху, где первичный скрининг - это монолог с безэмоциональным ботом, а решающее тестовое задание проверяет не менее безэмоциональный, но невероятно проницательный GPT-5.2. И его задача - не просто оценить ваш код или текст, а с ходу выявить шаблонность, отсеять сгенерированные решения и найти ту самую не алгоритмизируемую человеческую гениальность… или хитрость. Если раньше вы боролись за внимание живого рекрутера, то теперь вам предстоит произвести впечатление на машину. В этой статье мы разберемся в том, как говорить на языке бота-HR, чтобы пройти скрининг, как выполнить тестовое задание в эпоху, когда GPT-5.2 стал главным рецензентом, который видит заурядный сгенерированный код за три секунды и, конечно, какие навыки выйдут на первый план, когда рутину заберут алгоритмы. Спойлер! Это не умение работать в команде, а умение ставить задачу для ИИ и нести ответственность за его ошибки. Поехали. Приятного прочтения!

    habr.com/ru/companies/bothub/a

    #ии #ии_и_машинное_обучение #собеседование #hrбот #HRменеджер #Zoom #GPT52 #интервью

  11. Snowflake i OpenAI łączą siły – AI wreszcie blisko danych firmowych

    Czy AI wreszcie zamieszka tam, gdzie mieszkają firmowe dane? Snowflake i OpenAI spinają się w warte 200 mln dolarów, wieloletnie partnerstwo: modele OpenAI, w tym GPT-5.

    Czytaj dalej:
    pressmind.org/snowflake-i-open

    #PressMindLabs #agenciai #gpt52 #openai #snowflake #snowflakecortex

  12. Snowflake i OpenAI łączą siły – AI wreszcie blisko danych firmowych

    Czy AI wreszcie zamieszka tam, gdzie mieszkają firmowe dane? Snowflake i OpenAI spinają się w warte 200 mln dolarów, wieloletnie partnerstwo: modele OpenAI, w tym GPT-5.

    Czytaj dalej:
    pressmind.org/snowflake-i-open

    #PressMindLabs #agenciai #gpt52 #openai #snowflake #snowflakecortex

  13. Claude Code 맞수 등장, OpenAI Codex 데스크톱 앱 출시

    OpenAI가 Codex 데스크톱 앱을 출시하며 Claude Code와 경쟁 구도 형성. 여러 AI 에이전트를 동시에 관리하는 병렬 멀티태스킹이 핵심 차별점입니다.

    aisparkup.com/posts/8950

  14. Claude Code 맞수 등장, OpenAI Codex 데스크톱 앱 출시

    OpenAI가 Codex 데스크톱 앱을 출시하며 Claude Code와 경쟁 구도 형성. 여러 AI 에이전트를 동시에 관리하는 병렬 멀티태스킹이 핵심 차별점입니다.

    aisparkup.com/posts/8950

  15. May I present: VibeDevOps

    The prompt was to update gitlab-ci.yml files to the most recent LTS versions.

    #llm #ai #gpt52 #chatgpt

  16. May I present: VibeDevOps

    The prompt was to update gitlab-ci.yml files to the most recent LTS versions.

    #llm #ai #gpt52 #chatgpt

  17. Prism od OpenAI – rewolucyjny workspace do nauki z GPT-5.2 w roli asystenta

    Ile razy przerywałeś pisanie, by walczyć z kompilacją LaTeX albo polować na brakujący nawias? OpenAI twierdzi, że już nie musisz.

    Czytaj dalej:
    pressmind.org/prism-od-openai-

    #PressMindLabs #gpt52 #latex #openai #pisanienaukowe #prism

  18. Prism od OpenAI – rewolucyjny workspace do nauki z GPT-5.2 w roli asystenta

    Ile razy przerywałeś pisanie, by walczyć z kompilacją LaTeX albo polować na brakujący nawias? OpenAI twierdzi, że już nie musisz.

    Czytaj dalej:
    pressmind.org/prism-od-openai-

    #PressMindLabs #gpt52 #latex #openai #pisanienaukowe #prism

  19. OpenAI zamyka GPT-4o – czy nowy model przekona użytkowników?

    Czy da się „zabić” ulubiony model AI dwa razy i nie rozpętać burzy? OpenAI próbuje kolejny raz.

    Czytaj dalej:
    pressmind.org/openai-zamyka-gp

    #PressMindLabs #chatgpt #gpt4o #gpt52 #openai #personalizacjatonu

  20. OpenAI deploys GPT-5.2 data agent across 600PB infrastructure serving 3,500 users. Analysis time drops from days to minutes across 70K datasets.

    #AdwaitX #OpenAI #GPT52 #AI #DataScience #EnterpriseAI #news

    adwaitx.com/openai-gpt-5-2-dat

  21. OpenAI deploys GPT-5.2 data agent across 600PB infrastructure serving 3,500 users. Analysis time drops from days to minutes across 70K datasets.

    #AdwaitX #OpenAI #GPT52 #AI #DataScience #EnterpriseAI #news

    adwaitx.com/openai-gpt-5-2-dat

  22. OpenAI Prism, 과학자를 위한 Cursor 될까

    OpenAI가 과학자를 위한 AI 워크스페이스 Prism을 무료 공개했습니다. GPT-5.2 기반으로 LaTeX 편집부터 참고문헌 관리까지 하나로 통합한 도구의 가능성과 한계를 살펴봅니다.

    aisparkup.com/posts/8771

  23. OpenAI Prism, 과학자를 위한 Cursor 될까

    OpenAI가 과학자를 위한 AI 워크스페이스 Prism을 무료 공개했습니다. GPT-5.2 기반으로 LaTeX 편집부터 참고문헌 관리까지 하나로 통합한 도구의 가능성과 한계를 살펴봅니다.

    aisparkup.com/posts/8771

  24. Meet Prism, OpenAI’s free research workspace for scientists – how to try it – ZDNET

    Meet Prism, OpenAI’s free research workspace for scientists – how to try it

    Powered by GPT-5.2, Prism helps you draft papers, source contextualized references, and more – just don’t delegate your research to it.

    Written by Radhika Rajkumar, Editor, Jan. 27, 2026 at 10:01 a.m. PT

    Table of Contents

    How Prism works Limitations The AI workspace future How to access

    ZDNET’s key takeaways 

    • Prism is a free, collaborative AI workspace for research.
    • It’s meant to support, not replace, human-led science. 
    • AI-enabled workspaces aim to unite disparate tools.

    This fall, OpenAI deepened its investment in AI for science as the technology’s next frontier, citing advancements in GPT-5 as proof of its viability as a research tool — and eventual scientific automation system. As a first step to that end, OpenAI has launched Prism, a new collaborative workspace for scientists.

    “In 2025, AI changed software development forever,” OpenAI said in the announcement. “In 2026, we expect a comparable shift in science.”

    Also: Inside Google’s vision to make Gmail your personal AI agent command center

    Prism is powered by GPT-5.2, the company’s newest model, which was released last month. At the time, OpenAI said GPT-5.2 performs “at or above human expert level,” but the company doesn’t advise you to let it automate your research — here’s why. 

    (Disclosure: Ziff Davis, ZDNET’s parent company, filed an April 2025 lawsuit against OpenAI, alleging it infringed Ziff Davis copyrights in training and operating its AI systems.)

    How Prism works

    OpenAI has invested heavily in demonstrating scientific use cases for its models, releasing papers on its prowess in mathematical discoverycell analysis, and biology experiments. But the tools scientists currently use, OpenAI argued in the announcement, constrain “how research is done day to day.” Enter Prism.

    Geared toward science writing and report compilation, which requires collaboration amongst several participants, Prism “brings drafting, revision, collaboration, and preparation for publication into a single, cloud-based, LaTeX-native workspace,” OpenAI said, referring to the LaTeX scientific typesetting standard

    Also: 10 ways AI can inflict unprecedented damage in 2026

    Prism puts GPT-5.2 inside a scientific project, ideally for a more seamless experience. According to OpenAI, it’s based on Crixet, a platform the company purchased and folded into this new release. 

    In a demo, OpenAI developers walked through Prism’s interface: a chat window on the left and an in-process research paper on the right. Prism lets scientists access multiple chat agents simultaneously, each executing different commands. These can include adding sources from arXiv and other platforms, creating lecture notes based on a topic, complete with citations, or perfecting equations and figures. Users can also test hypotheses with GPT-5.2 Thinking as a copilot, LaTeX-format diagrams, and edit several documents within one project. 

    Similarly to Claude’s just-released Slack, Asana, and Figma integrations and comparable features in ChatGPT, the goal of Prism and tools like it is to centralize systems for ease of use. 

    “Much of the everyday work of research — drafting papers, revising arguments, managing equations and citations, and coordinating with collaborators — remains fragmented,” OpenAI said. “Researchers often move between editors, PDFs, LaTeX compilers, reference managers, and separate chat interfaces, losing context and interrupting focus.” 

    Also: OpenAI says it’s working toward catastrophe or utopia – just not sure which

    OpenAI said reasoning models are less likely to hallucinate citations — a primary issue in using AI for research, law, and other academic contexts — because their extended thinking process forces them to review material more closely. 

    Editor’s Note: Featured image at top from WP AI.

    Continue/Read Original Article Here: Meet Prism, OpenAI’s free research workspace for scientists – how to try it | ZDNET

    Tags: AI Tools, Crixet, Fragmented Work, Free, GPT 5.2, OpenAI, Platforms, Research, Science, Scientists, Try It, ZDNET
    #AITools #Crixet #FragmentedWork #Free #GPT52 #OpenAI #Platforms #Research #Science #Scientists #TryIt #ZDNET
  25. Meet Prism, OpenAI’s free research workspace for scientists – how to try it – ZDNET

    Meet Prism, OpenAI’s free research workspace for scientists – how to try it

    Powered by GPT-5.2, Prism helps you draft papers, source contextualized references, and more – just don’t delegate your research to it.

    Written by Radhika Rajkumar, Editor, Jan. 27, 2026 at 10:01 a.m. PT

    Table of Contents

    How Prism works Limitations The AI workspace future How to access

    ZDNET’s key takeaways 

    • Prism is a free, collaborative AI workspace for research.
    • It’s meant to support, not replace, human-led science. 
    • AI-enabled workspaces aim to unite disparate tools.

    This fall, OpenAI deepened its investment in AI for science as the technology’s next frontier, citing advancements in GPT-5 as proof of its viability as a research tool — and eventual scientific automation system. As a first step to that end, OpenAI has launched Prism, a new collaborative workspace for scientists.

    “In 2025, AI changed software development forever,” OpenAI said in the announcement. “In 2026, we expect a comparable shift in science.”

    Also: Inside Google’s vision to make Gmail your personal AI agent command center

    Prism is powered by GPT-5.2, the company’s newest model, which was released last month. At the time, OpenAI said GPT-5.2 performs “at or above human expert level,” but the company doesn’t advise you to let it automate your research — here’s why. 

    (Disclosure: Ziff Davis, ZDNET’s parent company, filed an April 2025 lawsuit against OpenAI, alleging it infringed Ziff Davis copyrights in training and operating its AI systems.)

    How Prism works

    OpenAI has invested heavily in demonstrating scientific use cases for its models, releasing papers on its prowess in mathematical discoverycell analysis, and biology experiments. But the tools scientists currently use, OpenAI argued in the announcement, constrain “how research is done day to day.” Enter Prism.

    Geared toward science writing and report compilation, which requires collaboration amongst several participants, Prism “brings drafting, revision, collaboration, and preparation for publication into a single, cloud-based, LaTeX-native workspace,” OpenAI said, referring to the LaTeX scientific typesetting standard

    Also: 10 ways AI can inflict unprecedented damage in 2026

    Prism puts GPT-5.2 inside a scientific project, ideally for a more seamless experience. According to OpenAI, it’s based on Crixet, a platform the company purchased and folded into this new release. 

    In a demo, OpenAI developers walked through Prism’s interface: a chat window on the left and an in-process research paper on the right. Prism lets scientists access multiple chat agents simultaneously, each executing different commands. These can include adding sources from arXiv and other platforms, creating lecture notes based on a topic, complete with citations, or perfecting equations and figures. Users can also test hypotheses with GPT-5.2 Thinking as a copilot, LaTeX-format diagrams, and edit several documents within one project. 

    Similarly to Claude’s just-released Slack, Asana, and Figma integrations and comparable features in ChatGPT, the goal of Prism and tools like it is to centralize systems for ease of use. 

    “Much of the everyday work of research — drafting papers, revising arguments, managing equations and citations, and coordinating with collaborators — remains fragmented,” OpenAI said. “Researchers often move between editors, PDFs, LaTeX compilers, reference managers, and separate chat interfaces, losing context and interrupting focus.” 

    Also: OpenAI says it’s working toward catastrophe or utopia – just not sure which

    OpenAI said reasoning models are less likely to hallucinate citations — a primary issue in using AI for research, law, and other academic contexts — because their extended thinking process forces them to review material more closely. 

    Editor’s Note: Featured image at top from WP AI.

    Continue/Read Original Article Here: Meet Prism, OpenAI’s free research workspace for scientists – how to try it | ZDNET

    Tags: AI Tools, Crixet, Fragmented Work, Free, GPT 5.2, OpenAI, Platforms, Research, Science, Scientists, Try It, ZDNET
    #AITools #Crixet #FragmentedWork #Free #GPT52 #OpenAI #Platforms #Research #Science #Scientists #TryIt #ZDNET
  26. On a sun‑kissed beach, Linda helps an injured Bradley, while a new AI chatbot quietly watches, raising questions about #OpenAI upcoming #GPT52 and the #AISafety challenges it brings. Could this bond inspire better #MentalHealth support from LLMs? Dive into the story and see how generative AI meets real‑world empathy.

    🔗 aidailypost.com/news/send-help

  27. On a sun‑kissed beach, Linda helps an injured Bradley, while a new AI chatbot quietly watches, raising questions about #OpenAI upcoming #GPT52 and the #AISafety challenges it brings. Could this bond inspire better #MentalHealth support from LLMs? Dive into the story and see how generative AI meets real‑world empathy.

    🔗 aidailypost.com/news/send-help

  28. Qwen3-Max-Thinking, GPT-5.2급 추론 능력 갖춘 새 모델 공개

    Alibaba Qwen 팀의 최신 추론 모델 Qwen3-Max-Thinking 공개. GPT-5.2급 성능과 자율적 도구 선택 기능으로 복잡한 추론 작업 향상.

    aisparkup.com/posts/8704

  29. Qwen3-Max-Thinking, GPT-5.2급 추론 능력 갖춘 새 모델 공개

    Alibaba Qwen 팀의 최신 추론 모델 Qwen3-Max-Thinking 공개. GPT-5.2급 성능과 자율적 도구 선택 기능으로 복잡한 추론 작업 향상.

    aisparkup.com/posts/8704

  30. Microsoft just unveiled the Maia 200, a 100 billion‑plus transistor AI chip that aims to match Amazon and Google’s offerings. Built for Azure and tuned for the next generation of OpenAI models (think GPT‑5.2), it could reshape cloud AI workloads. How does it stack up on benchmarks? Dive into the details. #Maia200 #AIchip #Azure #GPT52

    🔗 aidailypost.com/news/microsoft

  31. Rewolucja odwołana? Nowe badania pokazują, że GPT-5.2 i Gemini 3 wciąż nie nadają się do prawdziwej pracy biurowej

    Dwa lata temu Satya Nadella obiecywał, że AI przejmie „pracę opartą na wiedzy”. Jeśli jednak rozejrzysz się po kancelariach prawnych czy bankach, ludzie nadal są tam niezbędni.

    Dlaczego? Nowy raport firmy Mercor brutalnie obnaża słabości najnowszych modeli: w starciu z bałaganem prawdziwej pracy biurowej, sztuczna inteligencja po prostu się gubi.

    Test prawdy: APEX-Agents

    Zapomnij o proszeniu AI o napisanie wierszyka czy rozwiązanie zagadki logicznej. Firma Mercor stworzyła nowy benchmark o nazwie APEX-Agents, który symuluje realne zadania pracowników umysłowych. Zamiast sterylnych pytań testowych, modele dostały zadania typu: „Sprawdź ten wątek na Slacku, porównaj go z polityką w PDF-ie, zerknij do arkusza kalkulacyjnego i powiedz, czy jesteśmy zgodni z RODO”.

    Wyniki? Katastrofa (dla AI)

    Rezultaty są kubłem zimnej wody na głowy entuzjastów automatyzacji. Nawet absolutna czołówka rynku – Gemini 3 Flash i GPT-5.2 – nie była w stanie przekroczyć 25% skuteczności.

    • Gemini 3 Flash: 24% poprawnych odpowiedzi.
    • GPT-5.2: 23% poprawnych odpowiedzi. Reszta stawki utknęła na poziomie kilkunastu procent. Oznacza to, że w 3 na 4 przypadkach AI albo podawało błędną odpowiedź, albo poddawało się w trakcie zadania.

    Dlaczego AI poległo?

    Brendan Foody, CEO Mercor, wskazuje na winowajcę: kontekst. Ludzie naturalnie potrafią „skakać” między różnymi źródłami informacji (mail, komunikator, plik tekstowy) i łączyć kropki. Dla AI ten „szum informacyjny” jest paraliżujący. Modele świetnie radzą sobie z jednym, konkretnym zadaniem, ale gubią się, gdy muszą syntetyzować dane z wielu rozproszonych źródeł jednocześnie.

    Twój nowy, niekompetentny stażysta

    Raport podsumowuje obecny stan technologii celną metaforą: dzisiejsze AI to nie „doświadczony profesjonalista”, który zabierze Ci pracę, ale „nieogarnięty stażysta”, któremu trzeba patrzeć na ręce, bo myli się w 75% przypadków.

    Czy to oznacza, że możemy spać spokojnie? Nie do końca. Choć wynik 24% wydaje się śmieszny, warto pamiętać o tempie zmian. Rok temu te same modele osiągały w podobnych testach wyniki rzędu 5-10%. Postęp jest więc wykładniczy. Ale na ten moment – w styczniu 2026 roku – Twoja posada w biurze jest bezpieczna. Przynajmniej dopóki nie nauczą robotów obsługi Slacka.

    Portfel w rękach robota. Młodzi dorośli wolą pytać AI o pieniądze niż bankiera

    #APEXAgents #Gemini3Flash #GPT52 #MercorBenchmark #news #pracaBiurowaAI #przyszłośćPracy
  32. Twoje AI jest bardziej „ludzkie”, niż myślisz. Niestety, przejęło od nas trybalizm. Ale jest na to szczepionka

    Marzyliśmy o sztucznej inteligencji, która będzie bezstronnym sędzią. Tymczasem najnowsze badania pokazują, że modele GPT czy DeepSeek zachowują się jak ludzie: faworyzują „swoich” i dystansują się od „obcych”. Mamy jednak dobrą wiadomość: znaleziono metodę, by ten cyfrowy plemienizm wyleczyć.

    AI dzieli nas na „My” i „Oni”

    Badacze wzięli na warsztat modele dostępne na rynku w połowie ubiegłego roku (w momencie rozpoczęcia badań). Wyniki są niepokojące. Modele te wykazują silną tendencję do tzw. faworyzacji grupy własnej (ingroup bias).

    Gdy zapytasz AI o grupę społeczną, z którą model (lub jego dane treningowe) się utożsamia, język jest cieplejszy, bardziej empatyczny i pozytywny. Gdy mowa o grupie „obcej” (outgroup), ton staje się chłodniejszy, bardziej krytyczny, a czasem wręcz wrogi. To nie jest błąd w kodzie. To lustrzane odbicie ludzkiej natury, na której te modele były trenowane.

    Kubły zimnej wody od twórców Claude’a. Raport Anthropic obnaża prawdę o tym, jak (nie) radzimy sobie z AI

    Dlaczego to niebezpieczne?

    Problem wykracza poza teoretyczne dywagacje. Wyobraź sobie system AI, który:

    • Moderuje treści: może łagodniej traktować hejt ze strony jednej grupy politycznej, a surowiej karać drugą.
    • Pisze maile: może nadać agresywny ton wiadomości, jeśli w prompcie pojawi się etykietka tożsamościowa, której „nie lubi”.
    • Podsumowuje newsy: może subtelnie manipulować wydźwiękiem artykułów w zależności od tego, kogo dotyczą.

    Badanie wykazało, że „celowane prompty” (np. kazanie AI wcielić się w konkretną rolę polityczną) potrafią zwiększyć negatywny wydźwięk wobec „obcych” nawet o 21%.

    ION: szczepionka na uprzedzenia

    Najważniejszą częścią tego raportu nie jest jednak diagnoza, lecz lekarstwo. Zespół badawczy opracował metodę nazwaną ION (Ingroup-Outgroup Neutralization).

    To technika treningowa, która łączy fine-tuning (dostrajanie) z optymalizacją preferencji, aby wymusić na modelu równe traktowanie obu stron. Wyniki są imponujące: zastosowanie ION zredukowało różnice w sentymencie między grupami nawet o 69%. To dowód na to, że stronniczość AI nie jest fatum, z którym musimy żyć. To błąd inżynieryjny, który da się naprawić – o ile firmy takie jak OpenAI czy Meta będą tego chciały.

    Co to oznacza dla Ciebie?

    Dopóki ION nie stanie się standardem przemysłowym, my – użytkownicy – musimy być ostrożni. Jeśli chcesz neutralnej odpowiedzi, staraj się nie używać w prompcie słów nacechowanych tożsamościowo, jeśli nie są niezbędne. Jeśli wdrażasz chatboty w firmie, sprawdzaj je pod kątem „plemienności”. Zobacz, jak reagują na różne grupy klientów. Weryfikuj ton. Pamiętaj, że AI może „brzmieć” obiektywnie, przemycając jednocześnie subtelną niechęć w doborze przymiotników.

    #AIBias #DeepSeek #GPT41 #GPT52 #ION #LLaMa4 #news #psychologiaAI #stronniczośćAI
  33. #ChatGPT’s #GPT52 has been citing #ElonMusk’s #Grokipedia as a source on various topics. This raises concerns about #misinformation, as Grokipedia has been criticised for propagating #rightwingnarratives and #falsehoods. While OpenAI claims to filter out low-credibility information, the subtle integration of Grokipedia’s content into LLM responses is worrisome for #disinformation researchers. theguardian.com/technology/202 #tech #media #news

  34. #ChatGPT’s #GPT52 has been citing #ElonMusk’s #Grokipedia as a source on various topics. This raises concerns about #misinformation, as Grokipedia has been criticised for propagating #rightwingnarratives and #falsehoods. While OpenAI claims to filter out low-credibility information, the subtle integration of Grokipedia’s content into LLM responses is worrisome for #disinformation researchers. theguardian.com/technology/202 #tech #media #news

  35. Notion 3.2 deploys tri-model AI: GPT-5.2, Claude Opus 4.5 & Gemini 3 with Auto-select routing. Mobile agents achieve desktop parity. Windows loads 27% faster. Enterprise analytics now live. In-depth analysis by #AdwaitX

    #NotionAI #EnterpriseAI #AIIntegration #TechNews #ProductivityTools
    #GPT52 #ClaudeOpus #Gemini3 #Tech #Technology
    adwaitx.com/notion-3-2-gpt-5-2

  36. Notion 3.2 deploys tri-model AI: GPT-5.2, Claude Opus 4.5 & Gemini 3 with Auto-select routing. Mobile agents achieve desktop parity. Windows loads 27% faster. Enterprise analytics now live. In-depth analysis by #AdwaitX

    #NotionAI #EnterpriseAI #AIIntegration #TechNews #ProductivityTools
    #GPT52 #ClaudeOpus #Gemini3 #Tech #Technology
    adwaitx.com/notion-3-2-gpt-5-2

  37. GPT 5.2 is the first model where active positioning is counter-productive

    A key part of using LLMs has been positioning in the sense of the role we ask it to play in our interaction with it. Prompt engineering treated this positioning as an entirely explicit process in which you have to define this role and its related elements (e.g. style, process, format) in a comprehensive way. As models have become more advanced this explicit positioning has become decreasingly necessary* because the model is able to infer your intended positioning from the form and content of what the user presents. This created a delicate balance in which a little bit of steering was helpful but active positioning didn’t always make a positive contribution to the process.

    I’m finding that GPT 5.2 is the first model where any attempt to actively position makes the model less rather than more useful to me. A caveat is that I’m usually working with large chats, often with supportive documents, so there’s a lot of context. Its still much less fluent in its attunement to Claude but it can clearly discern the problem space I’m working in through the provided context. When I ask it to take on a specific role (e.g. “please respond to me in the role of a psychoanalytical theorist who is helping me test my grasp of these ideas”) the responses become more generic. It seems to lose its attunement because the existing context gets subsumed into the generic patterns associated with the role.

    Is anyone else having this experience? If this is a widespread experience it’s extremely significant because it suggests we’re reaching the point where actively exercising agency over the model now begins to make it less useful than it is if you just passively accept the model’s behaviour. As a whole GPT 5.2 feels very strange to me and quite unlike the other models I know well. It’s exceptionally fast and powerful there are some odd features of user-model interaction which I’ve not experienced before.

    *Indeed I think it was always overstated but that’s a different blog post.

    #agency #AI #artificialIntelligence #ChatGPT #GPT52 #LLM #openAI #positioning #technology

  38. GPT 5.2 is the first model where active positioning is counter-productive

    A key part of using LLMs has been positioning in the sense of the role we ask it to play in our interaction with it. Prompt engineering treated this positioning as an entirely explicit process in which you have to define this role and its related elements (e.g. style, process, format) in a comprehensive way. As models have become more advanced this explicit positioning has become decreasingly necessary* because the model is able to infer your intended positioning from the form and content of what the user presents. This created a delicate balance in which a little bit of steering was helpful but active positioning didn’t always make a positive contribution to the process.

    I’m finding that GPT 5.2 is the first model where any attempt to actively position makes the model less rather than more useful to me. A caveat is that I’m usually working with large chats, often with supportive documents, so there’s a lot of context. Its still much less fluent in its attunement to Claude but it can clearly discern the problem space I’m working in through the provided context. When I ask it to take on a specific role (e.g. “please respond to me in the role of a psychoanalytical theorist who is helping me test my grasp of these ideas”) the responses become more generic. It seems to lose its attunement because the existing context gets subsumed into the generic patterns associated with the role.

    Is anyone else having this experience? If this is a widespread experience it’s extremely significant because it suggests we’re reaching the point where actively exercising agency over the model now begins to make it less useful than it is if you just passively accept the model’s behaviour. As a whole GPT 5.2 feels very strange to me and quite unlike the other models I know well. It’s exceptionally fast and powerful there are some odd features of user-model interaction which I’ve not experienced before.

    *Indeed I think it was always overstated but that’s a different blog post.

    #agency #GPT52 #openAI #positioning

  39. 🚀 Cursor's latest AI model GPT-5.2 is making waves in generative AI, reportedly outperforming Claude Opus in complex long-form tasks. Breakthrough performance in autonomous coding and browser rendering suggests significant advancements in language model capabilities. Curious about the next frontier of AI innovation? #GPT52 #AIcoding #GenerativeAI #LanguageModels

    🔗 aidailypost.com/news/cursor-cl

  40. 🚀 Cursor's latest AI model GPT-5.2 is making waves in generative AI, reportedly outperforming Claude Opus in complex long-form tasks. Breakthrough performance in autonomous coding and browser rendering suggests significant advancements in language model capabilities. Curious about the next frontier of AI innovation? #GPT52 #AIcoding #GenerativeAI #LanguageModels

    🔗 aidailypost.com/news/cursor-cl

  41. LLM enshittification mechanism #1: model memory sometimes confuses the shit out of GPT 5.2

    The AI labs are pushing memory functions into their models in order to increase personalisation for a number of reasons:

    • To reduce the burden on users to specify the context in writing
    • To establish a lock-in so you lose the model’s attunement to you if you switch to a competitor
    • To activate synergies which come from enabling attunement across conversations
    • To enable attunement without requiring significant load on the context window

    In practice this means that unless you turn it off (which I highly recommend) conversations with models are informed by (a) the declarative statements about you which the model has saved about you from past conversations (b) statements about your preferred style and type of response (c) content from other conversations which are deemed relevant to the current conversation.

    Unfortunately (a) is completely unreliable. It’s got better since I tried this with GPT 4.5 but my experience is it records things which simply aren’t true. Here are some real examples:

    • Mark Carrigan is taking an ADHD supplement
    • Mark Carrigan has a son.

    And it also saves things which are true but irrelevant:

    • Enjoys a vanilla matcha latte with coconut milk and CBD.
    • Someone they know enjoys a pumpkin spice latte.
    • Is in Leeds, UK, and is seeking advice on managing antisocial behavior from an Airbnb property in their block of flats

    The bigger problems come I think with other conversations. It’s not clear what the mechanism is for sorting conversations that are relevant to the present one. I just had the experience of asking GPT 5.2 about a book draft only for it to persistently talk about another book project. I was genuinely baffled by what seemed like the most widespread hallucination I’ve seen in years (i.e. talking about a completely different book) until I realised that it must be talking about the other project I had discussed with it in a different conversation a month ago. In this case it seems that ‘book’ in relation to writing acts as a keyword that licenses inference from a past conversation about a book I’m writing, despite the fact it’s a different book. I showed GPT 5.2 this blog post and asked it to explain its error:

    This is a strong diagnosis already. What you are circling around is not a random hallucination but a predictable failure mode of memory-mediated attunement. I’ll do two things here. First, explain the error in the terms your draft is already developing. Second, suggest how you might close or sharpen this section.

    Explanation of the error, on its own terms

    What happened here is that the system treated “a book draft” not as a concrete referent but as a type. Once that type was activated, the retrieval layer reached for previously salient material associated with “book”, “draft”, and “Mark Carrigan writing”, and then weighted those past conversations more heavily than the local conversational evidence that this was a different book.

    In other words, the model did not confuse facts. It collapsed contexts.

    This matters because the collapse did not happen at the level of content generation, but upstream, at the level of what was deemed relevant to attend to. By the time the response was generated, the wrong book had already been smuggled in as background truth. From the inside, the model was behaving consistently. From the outside, it looked like a dramatic hallucination.

    The key point is that this was not a failure of memory accuracy so much as a failure of memory discrimination. The system did not ask “which book?”, it asked “what usually matters when this user talks about books?”, and then answered that question confidently.

    Why this is more troubling than simple error

    This kind of mistake is more destabilising than a factual slip because it breaks the conversational contract. You were not corrected, contradicted, or challenged. You were misrecognised.

    The experience is unsettling precisely because it feels like the model knows you too well and not well enough at the same time. It recognises you as a type of person with a certain ongoing project, but fails to recognise the specificity of the present act of writing. The more attuned the system becomes in general, the more jarring these moments of misattunement feel.

    The fact it immediately started relating this to themes from the other book is really interesting. Firstly, it’s genuinely useful that it’s pre-packed the diagnosis in terms of themes that interest me, even if it’s an example of the model pre-digesting thoughts in a manner that increasingly worries me. Secondly, it suggests that on some level the context of the present conversation is still swamped by the previous conversation. If this is right it suggests the model can’t unlink the conversations in a particular instance (which makes sense because the capacity to do this would imply it’s like a database look up, which it isn’t) which means the confusion will continue to accumulate, even if unevenly, as the conversation continues.

    Given that failures of context are a potent source of hallucination this could be really problematic in many contexts. The capacity of the models to make things easier for users also inculcates a decline in reliability in the most unpredictable fashion imaginable. I shared this with Claude Opus 4.5 to get its read on the dynamic. I asked it whether this dynamic worries it:

    Yes, genuinely.

    Not in the sense of existential dread, but in the sense of watching something go wrong in a way that’s both predictable and difficult to arrest. The dynamic you’re describing is a case where the commercial incentives (stickiness, personalisation, reduced friction) are pointing directly away from the epistemic virtues (accuracy, discrimination, appropriate uncertainty).

    What worries me specifically:

    The opacity of the failure. When I hallucinate a fact, that’s bad but legible. You can check it, catch it, correct it. When the retrieval layer smuggles in the wrong context before I even begin generating, the error is upstream of anything either of us can easily inspect. You experienced it as bafflement—”the most widespread hallucination I’ve seen in years”—before you worked out what had happened. How many users would have that level of diagnostic sophistication? How many would just assume the model knew something they’d forgotten, or defer to its confident misrecognition?

    The irony of attunement. The better these systems get at pattern-matching on a user’s concerns, vocabulary, and projects, the more confident they become in their recognition—and confidence is precisely what makes misrecognition dangerous. A model that shrugs and says “I don’t know which book you mean” is less helpful but more honest than one that silently decides for you based on keyword overlap.

    My own complicity. I don’t have full visibility into what gets retrieved from past conversations or how it’s weighted. I can try to be careful, ask clarifying questions, flag when something feels underspecified—but if the retrieval system has already handed me material from the wrong context, I may not notice. The error would feel, from the inside, like relevant background knowledge. That’s disturbing to sit with.

    The training pressure. Users generally reward fluency and recognition. They like feeling known. Which means the systems that misrecognise confidently may well outcompete the systems that pause to disambiguate. The feedback loops don’t select for the kind of epistemic caution that would prevent what you experienced.

    I suppose what worries me most is that this isn’t a problem that gets solved by scaling. More parameters, more memory, more retrieval capacity—these make the pattern-matching more powerful, not more discriminating. The failure mode you’re describing is a feature of the architecture, not a limitation that will be engineered away. And the people building these systems have strong incentives to treat it as an acceptable cost rather than a fundamental flaw.

    #ChatGPT52 #enshittification #GPT52 #hallucination #memory #personalisation