home.social

#summarization — Public Fediverse posts

Live and recent posts from across the Fediverse tagged #summarization, aggregated by home.social.

fetched live
  1. Эволюция подхода к сжатию контекста в AI Агентах

    AI-агенты - это цикл обмена сообщениями между пользователем и языковой моделью (LLM), где для ответа пользователю модель может обратиться к доступным ей напрямую инструментам (tools) или через настроенные для неё MCP. Каждое такое взаимодействие дописывает историю диалога. Для простоты часто считают, что вся эта история и уходит в любой следующий вызов модели - иначе она «не вспомнит», о чём шла речь, и криво соберёт ответ. История «не резиновая» - у современных моделей контекстное окно может быть огромным, но умение работать с длинным логом сильно зависит от того, насколько он структурирован. Плюс накопленная история разговора и реальный контекст, который модель видит на очередном ходе, не всегда одно и то же. В последнее время это один из фокусов развития агентских систем в рамках context engineering: что сжать, что оставить снаружи, что подтянуть инструментами только когда нужно. В этой статье хочу рассказать про эволюцию подхода работы с длинной историей и где мы находимся сейчас. Примеры будут на LangChain - на открытом стеке легко посмотреть реализацию, которая в готовых продуктах часто спрятана. При этом LangChain достаточно популярный, развивающийся фреймворк - остальные либо делают похожие вещи, либо сами опираются на него как на базу.

    habr.com/ru/articles/1069780/

    #агенты #машинное_обучение #ииагенты #context_engineering #summarization

  2. For the #ttrpg bubble

    Oh, did I even tell you that I've put the scripts I'm using for my TranscriptOMatic #roleplaying session transcription proof-of-concept into a Git repository?

    codeberg.org/Felicea/Transcrip

    Documentation of the live-transcription TranscriptOMatic part: info.zusammenkunft.net/shelves

    Documentation for the post-production is still in the making.

    #RPG #session #transcription #summarization #OpenSource #LocalLLM

  3. For the #ttrpg bubble

    Oh, did I even tell you that I've put the scripts I'm using for my TranscriptOMatic #roleplaying session transcription proof-of-concept into a Git repository?

    codeberg.org/Felicea/Transcrip

    Documentation of the live-transcription TranscriptOMatic part: info.zusammenkunft.net/shelves

    Documentation for the post-production is still in the making.

    #RPG #session #transcription #summarization #OpenSource #LocalLLM

  4. 🎉 Oh, look! Another #GitHub project's here to save us from the grueling task of downloading and summarizing videos - because reading is so 2022! 📼🙄 #OpenBrief promises to be "local-first," which is just a fancy way of saying, "Good luck figuring this out without WiFi!" 😂🔌
    github.com/tantara/openbrief #video #summarization #local-first #tech #humor #HackerNews #ngated

  5. 🎉 Oh, look! Another #GitHub project's here to save us from the grueling task of downloading and summarizing videos - because reading is so 2022! 📼🙄 #OpenBrief promises to be "local-first," which is just a fancy way of saying, "Good luck figuring this out without WiFi!" 😂🔌
    github.com/tantara/openbrief #video #summarization #local-first #tech #humor #HackerNews #ngated

  6. We just dropped a new #tutorial 🙂
    youtu.be/VKqWNHagZks

    This time we’re looking at how to extract all the text from a #webpage (or just a specific section).

    It’s a simple node, but it’s actually the starting point for a lot of #workflows (like AI #summarization .

    Still early days for the project, so feedback really helps 🙏

    awflow.io

  7. #Text #Summarization is the process of distilling a large document into a concise version while preserving its core meaning and factual integrity.

    By utilizing Natural Language Processing (NLP), it helps users quickly digest vast amounts of data, such as news articles, legal papers, or research reports.

    knowledgezone.co.in/trends/bro

  8. #Text #Summarization is the process of distilling a large document into a concise version while preserving its core meaning and factual integrity.

    By utilizing Natural Language Processing (NLP), it helps users quickly digest vast amounts of data, such as news articles, legal papers, or research reports.

    knowledgezone.co.in/trends/bro

  9. Hi All Mastodonians,
    I sometimes won't checkout mastodon for 3 to 4 days or even a week. How do you all catchup with things. I tried adding some accounts to a news list. But still they are hard to catch up since there are too many posts.
    Anyone using any AI or something to summarise the posts? Any clients doing that ? 🤔

    #mastodon #feeds #toots #CatchingUp #content #summarization #fediverse #fedi #posts #ai

  10. Как я сделал автоматический Телеграм канал с помощью Gmail и OpenAI API

    Как мы сделали автоматический Телеграм канал который по апи собирает новостные рассылки, суммаризирует и постит в Телеграм.

    habr.com/ru/articles/929108/

    #newsletters #summarization #telegram

  11. 🤖 Resource-Efficient & Effective Code Summarization

    (funny, it's harder to make an AI tell what code does than it is to make one write code...)

    arxiv.org/abs/2502.03617

    #ai #llm #ml #summarization #coding #programming #reverseengineering

  12. 🤖 Resource-Efficient & Effective Code Summarization

    (funny, it's harder to make an AI tell what code does than it is to make one write code...)

    arxiv.org/abs/2502.03617

    #ai #llm #ml #summarization #coding #programming #reverseengineering

  13. 📚 **AI-Powered Study Guide Summarization!** 🤖✨

    Struggling with **long study materials**? Learn how **Prompt Engineering** can help you craft AI queries that generate **concise, structured, and effective summaries**! 🧠🔍

    📖 promptengineering.ninja/p/prom

    #AI #PromptEngineering #StudySmart #EdTech #MachineLearning #ArtificialIntelligence #Summarization #AIinEducation

  14. → We’re Doing What Searchbots Can’t
    thewalrus.ca/were-doing-what-s

    “[Summarization tools incorporated into the search engines are] a boon for people seeking quick answers, but a bane for publishers. Disincentivizing curious users from clicking through to a news site for additional information—a trend called zero-click search—sends less traffic to media outlets that invest in the costly #reporting that #AI machines are scraping, strip-mining, and synthesizing.”

    #Summarization #search #news #media

  15. → We’re Doing What Searchbots Can’t
    thewalrus.ca/were-doing-what-s

    “[Summarization tools incorporated into the search engines are] a boon for people seeking quick answers, but a bane for publishers. Disincentivizing curious users from clicking through to a news site for additional information—a trend called zero-click search—sends less traffic to media outlets that invest in the costly #reporting that #AI machines are scraping, strip-mining, and synthesizing.”

    #Summarization #search #news #media

  16. #Summarization: "Australian Government Trial Finds #AI is Much Worse Than Humans at Summarizing" & More AI News Headlines ow.ly/fNho50TfUAu

  17. #Summarization: "Australian Government Trial Finds #AI is Much Worse Than Humans at Summarizing" & More AI News Headlines ow.ly/fNho50TfUAu

  18. Обзор приложения NotebookLM

    Приложение под названием NotebookLM ( notebooklm.google.com/ ) было выпущено компанией Google около года назад, и на Хабре было по этому поводу два кратких анонса в прошлом году ( раз , два ). На мой взгляд, оно заслуживает обзора чуть более подробного чем эти краткие сообщения, так что попробую восполнить этот пробел. NotebookLM - это инструмент на основе ИИ, который позволяет относительно быстро, удобно и без лишних телодвижений получить краткий разносторонний обзор (саммари) объемных документов (книг, статей), а также интерактивно взаимодействовать с ними (задавать вопросы, касающиеся их содержания). В моем понимании он представляет собой надстройку над "обычным ИИ-чатом", которому в контекст загружен интересующий пользователя документ. Эта надстройка включает в себя: 1. Набор из нескольких преднастроенных стандартизованных промптов, доступных в один клик и ориентированных на работу с объемными текстами ("Составь мне оглавление", "Составь мне FAQ на основе этого текста", и т.п.) 2. Интерфейсное решение ("карточки-плитки на рабочем столе"), которое по замыслу разработчиков, видимо, должно быть более удобным чем "обычный (линейный) чат" 3. Интерфейс чата, который при взаимодействии с текстом в формате "вопрос-ответ" отображает не только ответы на задаваемые вопросы, но и фрагменты соответствующего исходного текста, а также ссылки на конкретные параграфы полного текста-источника. Посмотрим как это работает

    habr.com/ru/articles/839668/

    #google #summarization #notebooklm #продуктивность

  19. “Accessible documents result in a more accurate AI summarization of each document.”

    Use this as motivation to improve accessibility leveraging corporate interest in AI.
    It can also be used as an AI use-case, to improve accessibility of documents using AI.

    Use this as motivation to improve the data classification of documents.
    It can also be used as an AI use case.

    #AI #accessibility #data #classification #summarization #infosec

  20. Как анализировать тысячи отзывов с ChatGPT? Частые ошибки и пример на реальных данных

    В этой статье я расскажу про свой опыт решения рабочей задачи — анализ отзывов о компании от пользователей. Мы разберем возможные ошибки и посмотрим на пример кода и реальных данных. Гайд будет полезен всем, у кого нет большого опыта в анализе данных или работе с LLM через API.

    habr.com/ru/articles/821287/

    #llm #gpt #chatgpt #python #clustering #kmeans #tsne #visualization #summarization #data_analysis

  21. Автоматизируем поиск ценной информации в групповых чатах Telegram с помощью LLM

    Устали мониторить бесконечные групповые чаты в Telegram в поисках важной информации? Решение есть! Пишем компактное приложение на Python, которое будет делать это за нас с использованием LLM.

    habr.com/ru/articles/804111/

    #telegram #chatgpt #llm #summarization #автоматизация #боты #gpt #python

  22. As the #GEM team already mentioned, we have endorsed the #data2text and #summarization shared tasks taking place this year: gem-benchmark.com/shared_task

    For data-to-text, there are two different datasets and you can choose to work with factual, counterfactual, or fictional versions of the datasets.

    For summarization, you can work on Swahili, cross-lingual summarizaion, or summarizing English book chapters.

    Interesting challenges with a deadline of 5 April with human evaluations starting on the 6th

  23. New efficient eval results

    1. A few examples are enough for Human preference to be clear, automatic metrics also don't need too many
    2. Context may change which model is preferred

    arxiv.org/abs/2402.18756
    #evaluation #nlp #nlproc #ML #summarization #efival

  24. New efficient eval results

    1. A few examples are enough for Human preference to be clear, automatic metrics also don't need too many
    2. Context may change which model is preferred

    arxiv.org/abs/2402.18756
    #evaluation #nlp #nlproc #ML #summarization #efival

  25. Предсказать ошибку. Как методы оценки неопределенности помогают повышать качество seq2seq-моделей

    Всем привет! Меня зовут Артём Важенцев , я аспирант в Сколтехе и младший научный сотрудник AIRI. Наша группа занимается исследованием и разработкой новых методов оценивания неопределенности для языковых моделей. Этим летом мы опубликовали две статьи на ACL 2023 . Про одну из них я уже рассказывал в одном из предыдущих текстов — там мы описали новый гибридный метод оценивания неопределенности для задачи выборочной классификации текстов. Другая же статья про то, как мы адаптировали современные методы оценивания неопределенности на основе скрытого представления модели для задачи генерации текста, а так же показали их высокое качество и скорость работы для задачи обнаружения примеров вне обучающего распределения. Ниже я хотел бы подробнее рассказать об используемых методах и результатах, которые мы получили.

    habr.com/ru/companies/airi/art

    #uncertainty_estimation #natural_language_processing #machine_translation #question_answering #summarization #seq2seq

  26. Yesterday at #TPDL2023 David Pride presented “CORE-GPT: Combining Open Access research and large language models for credible, trustworthy question answering”

    Rather than #ZeroShot question/answering, Pride’s team combines the #CORE #OpenAccess dataset with #ElasticSearch to create #FewShot prompts that leverage the strength of combining #search results with the #LLM’s (#GPT) #summarization abilities to produce an answer to a user’s question including citations.

    Ref: doi.org/10.1007/978-3-031-4384

  27. What about that metadata that is present? Grusky et al. (doi.org/10.18653/v1/N18-1065 ) realized that, because page authors create that metadata, it can serve as ground truth to evaluate #Automatic #Summarization.

    We analyzed pages from #WebArchiving and saw how this metadata evolved. By 2010 we saw a metadata explosion with the use of #Twitter Cards, Open Graph Protocol, #Facebook Tracking, and more. Things like Twitter cards created a metadata renaissance for HTML.

    Ref: doi.org/10.1109/JCDL52503.2021

  28. Social cards are generated based on #metadata present in web pages. If the author does not create the metadata, the service will not create the card. What do we do for web pages that predate this metadata?

    A lot of #Automatic #Summarization techniques can help us create the description part of social cards, but what about the image? In 2021, we found that Random Forest #MachineLearning can help choose the correct image using easy-to-calculate features.

    Ref: doi.org/10.1145/3447535.346250

  29. In 2020, we developed a special tool, MementoEmbed, for generating/extracting metadata from archived web pages. We presented this tool at the Web Archiving and Digital Libraries Workshop (WADL2020).

    We found out that #Twitter, #Facebook, #Tumblr, and others could not reliably create cards for archived web pages. We use MementoEmbed’s cards in #Storytelling with our tool Raintale to create a #Visualization of this #Summarization.

    #WebArchiving #DigitalPreservation

    Ref: arxiv.org/abs/2008.00137

  30. Elon is planning to effectively kill social cards on #Twitter. Social cards were a big part of my dissertation work. I published a few papers about generating them via #ComputerVision, #NLP, and #MachineLearning because they make for nice bits of document #Summarization and #Storytelling. Now Musk wants them gone to force journalists to write articles directly on Twitter.

    Ref (paywall): fortune.com/2023/08/21/elon-mu
    Ref (article about paywalled article): 9to5mac.com/2023/08/21/twitter

    #TwitterMigration

  31. Machine Learning: Translating languages, transcribing voices, and summarizing content in the blink of an eye! The future of seamless communication and information consumption is here. col.la/mlscript #machinelearning #transcription #translation #summarization

  32. Proposing 😄 SMEIL (Semantic Matrix Embeddings and Information Loss) scores for #summarization of text as an alternative to ROUGE, BERTScore and similar: Authors: Me, and #GPT4: chat.openai.com/share/bfd426f5

  33. @ingorohlfing pity. If I cannot trust a tool to provide accurate information it's useless. I was hoping when it's about a #summarization tasks these new #llm tools would be able and be constrained to not make things up. After all, well working #abstractiveSummarization models already exist for some time. See, e.g., paperswithcode.com/task/abstra

  34. Embrace Lifelong Learning, Or Else • TechNotes Blog

    For students, summarizing can be straightforward. They can learn to circle the main idea, underline details, and use sentence starters to summarize.

    #edtech #summarization #future #education #edutooter #tcea @edutooters

    blog.tcea.org/embrace-lifelong