home.social

#tfidf — Public Fediverse posts

Live and recent posts from across the Fediverse tagged #tfidf, aggregated by home.social.

fetched live
  1. How a computer reads text - from counting words to vectors

    From tokenization through TF-IDF and Markov chains, to Word2Vec. How a computer turns text into numb...

    gruszka.dev/en/how-computer-re
    #llm #ai #nlp #tokenization #word2vec #embeddings #tfidf #markov #bayes #languagemodels

  2. How a computer reads text - from counting words to vectors

    From tokenization through TF-IDF and Markov chains, to Word2Vec. How a computer turns text into numb...

    gruszka.dev/en/how-computer-re
    #llm #ai #nlp #tokenization #word2vec #embeddings #tfidf #markov #bayes #languagemodels

  3. How a computer reads text - from counting words to vectors

    From tokenization through TF-IDF and Markov chains, to Word2Vec. How a computer turns text into numb...

    gruszka.dev/en/how-computer-re
    #llm #ai #nlp #tokenization #word2vec #embeddings #tfidf #markov #bayes #languagemodels

  4. How a computer reads text - from counting words to vectors

    From tokenization through TF-IDF and Markov chains, to Word2Vec. How a computer turns text into numb...

    gruszka.dev/en/how-computer-re
    #llm #ai #nlp #tokenization #word2vec #embeddings #tfidf #markov #bayes #languagemodels

  5. How a computer reads text - from counting words to vectors

    From tokenization through TF-IDF and Markov chains, to Word2Vec. How a computer turns text into numb...

    gruszka.dev/en/how-computer-re
    #llm #ai #nlp #tokenization #word2vec #embeddings #tfidf #markov #bayes #languagemodels

  6. Jak komputer czyta tekst - od liczenia słów do wektorów

    Od tokenizacji przez TF-IDF i łańcuchy Markowa, aż po Word2Vec. Jak komputer zamienia tekst w liczby...

    gruszka.dev/jak-komputer-czyta
    #llm #ai #nlp #tokenizacja #word2vec #embeddings #tfidf #markow #bayes #languagemodels

  7. Jak komputer czyta tekst - od liczenia słów do wektorów

    Od tokenizacji przez TF-IDF i łańcuchy Markowa, aż po Word2Vec. Jak komputer zamienia tekst w liczby...

    gruszka.dev/jak-komputer-czyta
    #llm #ai #nlp #tokenizacja #word2vec #embeddings #tfidf #markow #bayes #languagemodels

  8. Jak komputer czyta tekst - od liczenia słów do wektorów

    Od tokenizacji przez TF-IDF i łańcuchy Markowa, aż po Word2Vec. Jak komputer zamienia tekst w liczby...

    gruszka.dev/jak-komputer-czyta
    #llm #ai #nlp #tokenizacja #word2vec #embeddings #tfidf #markow #bayes #languagemodels

  9. Jak komputer czyta tekst - od liczenia słów do wektorów

    Od tokenizacji przez TF-IDF i łańcuchy Markowa, aż po Word2Vec. Jak komputer zamienia tekst w liczby...

    gruszka.dev/jak-komputer-czyta
    #llm #ai #nlp #tokenizacja #word2vec #embeddings #tfidf #markow #bayes #languagemodels

  10. Jak komputer czyta tekst - od liczenia słów do wektorów

    Od tokenizacji przez TF-IDF i łańcuchy Markowa, aż po Word2Vec. Jak komputer zamienia tekst w liczby...

    gruszka.dev/jak-komputer-czyta
    #llm #ai #nlp #tokenizacja #word2vec #embeddings #tfidf #markow #bayes #languagemodels

  11. How a computer reads text - from counting words to vectors

    From tokenization through TF-IDF and Markov chains, to Word2Vec. How a computer turns text into numb...

    gruszka.dev/en/how-computer-re
    #llm #ai #nlp #tokenization #word2vec #embeddings #tfidf #markov #bayes #languagemodels

  12. How a computer reads text - from counting words to vectors

    From tokenization through TF-IDF and Markov chains, to Word2Vec. How a computer turns text into numb...

    gruszka.dev/en/how-computer-re
    #llm #ai #nlp #tokenization #word2vec #embeddings #tfidf #markov #bayes #languagemodels

  13. How a computer reads text - from counting words to vectors

    From tokenization through TF-IDF and Markov chains, to Word2Vec. How a computer turns text into numb...

    gruszka.dev/en/how-computer-re
    #llm #ai #nlp #tokenization #word2vec #embeddings #tfidf #markov #bayes #languagemodels

  14. От PDF к учебному модулю: практичный ML-пайплайн внутри LMS

    Всем привет, с вами Михаил Киселев, ML-разработчик в компании WebRise. И сегодня поговорим о практическом применении ML в образовании. Почему при горе регламентов, инструкций и методичек запуск нового курса всё равно растягивается на недели? И почему проблема часто не в LMS, а на шаг раньше — там, где знания в компании уже есть, а учебной структуры ещё нет?

    habr.com/ru/articles/1053450/

    #ml #lms #tfidf #kmeans #python #дистанционное_образование #дистанционное_обучение

  15. От PDF к учебному модулю: практичный ML-пайплайн внутри LMS

    Всем привет, с вами Михаил Киселев, ML-разработчик в компании WebRise. И сегодня поговорим о практическом применении ML в образовании. Почему при горе регламентов, инструкций и методичек запуск нового курса всё равно растягивается на недели? И почему проблема часто не в LMS, а на шаг раньше — там, где знания в компании уже есть, а учебной структуры ещё нет?

    habr.com/ru/articles/1053450/

    #ml #lms #tfidf #kmeans #python #дистанционное_образование #дистанционное_обучение

  16. От PDF к учебному модулю: практичный ML-пайплайн внутри LMS

    Всем привет, с вами Михаил Киселев, ML-разработчик в компании WebRise. И сегодня поговорим о практическом применении ML в образовании. Почему при горе регламентов, инструкций и методичек запуск нового курса всё равно растягивается на недели? И почему проблема часто не в LMS, а на шаг раньше — там, где знания в компании уже есть, а учебной структуры ещё нет?

    habr.com/ru/articles/1053450/

    #ml #lms #tfidf #kmeans #python #дистанционное_образование #дистанционное_обучение

  17. Jak komputer czyta tekst - od liczenia słów do wektorów

    Od tokenizacji przez TF-IDF i łańcuchy Markowa, aż po Word2Vec. Jak komputer zamienia tekst w liczby...

    gruszka.dev/jak-komputer-czyta
    #llm #ai #nlp #tokenizacja #word2vec #embeddings #tfidf #markow #bayes #languagemodels

  18. Jak komputer czyta tekst - od liczenia słów do wektorów

    Od tokenizacji przez TF-IDF i łańcuchy Markowa, aż po Word2Vec. Jak komputer zamienia tekst w liczby...

    gruszka.dev/jak-komputer-czyta
    #llm #ai #nlp #tokenizacja #word2vec #embeddings #tfidf #markow #bayes #languagemodels

  19. Jak komputer czyta tekst - od liczenia słów do wektorów

    Od tokenizacji przez TF-IDF i łańcuchy Markowa, aż po Word2Vec. Jak komputer zamienia tekst w liczby...

    gruszka.dev/jak-komputer-czyta
    #llm #ai #nlp #tokenizacja #word2vec #embeddings #tfidf #markow #bayes #languagemodels

  20. Jak komputer czyta tekst - od liczenia słów do wektorów

    Od tokenizacji przez TF-IDF i łańcuchy Markowa, aż po Word2Vec. Jak komputer zamienia tekst w liczby...

    gruszka.dev/jak-komputer-czyta
    #llm #ai #nlp #tokenizacja #word2vec #embeddings #tfidf #markow #bayes #languagemodels

  21. Jak komputer czyta tekst - od liczenia słów do wektorów

    Od tokenizacji przez TF-IDF i łańcuchy Markowa, aż po Word2Vec. Jak komputer zamienia tekst w liczby...

    gruszka.dev/jak-komputer-czyta
    #llm #ai #nlp #tokenizacja #word2vec #embeddings #tfidf #markow #bayes #languagemodels

  22. Эволюция 'More Like This'

    Во многих поисковых сценариях пользователь начинает не с пустой строки запроса, а с существующего результата. Пользователь открывает статью и хочет найти похожие материалы. Покупатель просматривает карточку товара и ищет близкие варианты. Инженер поддержки разбирает инцидент и хочет увидеть прошлые случаи с теми же симптомами. Во всех этих ситуациях у пользователя уже есть релевантный документ для начала поиска. Этот сценарий традиционно называют More Like This (MLT) : функцией поиска документов, похожих на выбранный. В статье под MLT понимается поиск от уже известного документа, а не от заново введённого запроса. Классический подход MLT (поиск похожих документов) основывался на сравнении текстовых совпадений. Современные реализации всё чаще используют эмбеддинги: числовые представления документов. Поисковый индекс хранит эмбеддинги в виде векторов, а поисковая система может находить документы с близкими векторными представлениями.

    habr.com/ru/articles/1042190/

    #nlp #обработка_естественного_языка #векторный_поиск #оптимизация_производительности #полнотекстовый_поиск #семантический_поиск #ранжирование_поиска #tfidf #bm25

  23. Эволюция 'More Like This'

    Во многих поисковых сценариях пользователь начинает не с пустой строки запроса, а с существующего результата. Пользователь открывает статью и хочет найти похожие материалы. Покупатель просматривает карточку товара и ищет близкие варианты. Инженер поддержки разбирает инцидент и хочет увидеть прошлые случаи с теми же симптомами. Во всех этих ситуациях у пользователя уже есть релевантный документ для начала поиска. Этот сценарий традиционно называют More Like This (MLT) : функцией поиска документов, похожих на выбранный. В статье под MLT понимается поиск от уже известного документа, а не от заново введённого запроса. Классический подход MLT (поиск похожих документов) основывался на сравнении текстовых совпадений. Современные реализации всё чаще используют эмбеддинги: числовые представления документов. Поисковый индекс хранит эмбеддинги в виде векторов, а поисковая система может находить документы с близкими векторными представлениями.

    habr.com/ru/articles/1042190/

    #nlp #обработка_естественного_языка #векторный_поиск #оптимизация_производительности #полнотекстовый_поиск #семантический_поиск #ранжирование_поиска #tfidf #bm25

  24. Эволюция 'More Like This'

    Во многих поисковых сценариях пользователь начинает не с пустой строки запроса, а с существующего результата. Пользователь открывает статью и хочет найти похожие материалы. Покупатель просматривает карточку товара и ищет близкие варианты. Инженер поддержки разбирает инцидент и хочет увидеть прошлые случаи с теми же симптомами. Во всех этих ситуациях у пользователя уже есть релевантный документ для начала поиска. Этот сценарий традиционно называют More Like This (MLT) : функцией поиска документов, похожих на выбранный. В статье под MLT понимается поиск от уже известного документа, а не от заново введённого запроса. Классический подход MLT (поиск похожих документов) основывался на сравнении текстовых совпадений. Современные реализации всё чаще используют эмбеддинги: числовые представления документов. Поисковый индекс хранит эмбеддинги в виде векторов, а поисковая система может находить документы с близкими векторными представлениями.

    habr.com/ru/articles/1042190/

    #nlp #обработка_естественного_языка #векторный_поиск #оптимизация_производительности #полнотекстовый_поиск #семантический_поиск #ранжирование_поиска #tfidf #bm25

  25. AI для PHP-разработчиков. Часть 6: Bag of Words и TF–IDF – как компьютер превращает текст в математику

    Когда мы говорим, что нейросети "понимают текст", легко забыть: компьютер изначально вообще не понимает слова. Для него текст – это набор чисел, статистики и векторов. В этой статье разберём Bag of Words и TF–IDF – фундаментальные подходы, с которых начинались NLP, поисковые системы и анализ текста. А заодно реализуем поиск похожих документов на чистом PHP без библиотек.

    habr.com/ru/articles/1034874/

    #php #machinelearning #bagofwords #tfidf #BoW #NLP #обработка_естественного_языка #cosine_similarity #векторизация_текста #машинное_обучение

  26. AI для PHP-разработчиков. Часть 6: Bag of Words и TF–IDF – как компьютер превращает текст в математику

    Когда мы говорим, что нейросети "понимают текст", легко забыть: компьютер изначально вообще не понимает слова. Для него текст – это набор чисел, статистики и векторов. В этой статье разберём Bag of Words и TF–IDF – фундаментальные подходы, с которых начинались NLP, поисковые системы и анализ текста. А заодно реализуем поиск похожих документов на чистом PHP без библиотек.

    habr.com/ru/articles/1034874/

    #php #machinelearning #bagofwords #tfidf #BoW #NLP #обработка_естественного_языка #cosine_similarity #векторизация_текста #машинное_обучение

  27. AI для PHP-разработчиков. Часть 6: Bag of Words и TF–IDF – как компьютер превращает текст в математику

    Когда мы говорим, что нейросети "понимают текст", легко забыть: компьютер изначально вообще не понимает слова. Для него текст – это набор чисел, статистики и векторов. В этой статье разберём Bag of Words и TF–IDF – фундаментальные подходы, с которых начинались NLP, поисковые системы и анализ текста. А заодно реализуем поиск похожих документов на чистом PHP без библиотек.

    habr.com/ru/articles/1034874/

    #php #machinelearning #bagofwords #tfidf #BoW #NLP #обработка_естественного_языка #cosine_similarity #векторизация_текста #машинное_обучение

  28. FYI: New Search Engine: Relevance Factors & Ranking Explained! #shorts: A new search engine must prioritize factors beyond simple text matches. Considerations include article freshness, popularity, and source to ensure the most relevant results. Location matters for restaurant searches, ensuring nearby options appear first. #searchengine #relevance #TFIDF #location #proximity youtube.com/shorts/nmWQGIm_hmc

  29. FYI: New Search Engine: Relevance Factors & Ranking Explained! #shorts: A new search engine must prioritize factors beyond simple text matches. Considerations include article freshness, popularity, and source to ensure the most relevant results. Location matters for restaurant searches, ensuring nearby options appear first. #searchengine #relevance #TFIDF #location #proximity youtube.com/shorts/nmWQGIm_hmc

  30. FYI: New Search Engine: Relevance Factors & Ranking Explained! #shorts: A new search engine must prioritize factors beyond simple text matches. Considerations include article freshness, popularity, and source to ensure the most relevant results. Location matters for restaurant searches, ensuring nearby options appear first. #searchengine #relevance #TFIDF #location #proximity youtube.com/shorts/nmWQGIm_hmc

  31. FYI: New Search Engine: Relevance Factors & Ranking Explained! #shorts: A new search engine must prioritize factors beyond simple text matches. Considerations include article freshness, popularity, and source to ensure the most relevant results. Location matters for restaurant searches, ensuring nearby options appear first. #searchengine #relevance #TFIDF #location #proximity youtube.com/shorts/nmWQGIm_hmc

  32. Почему Andrej Karpathy использует SVM в 2026 году (и вам тоже стоит)

    На arXiv каждый день публикуются сотни статей по машинному обучению. Читать всё — нереально, а пропустить что-то важное — обидно. Andrej Karpathy, бывший Director of AI в Tesla и соавтор курса Stanford CS231n, решил эту проблему неожиданным способом. Он выбрал не BERT, не GPT и не какой-нибудь модный трансформер. Он остановился на добром старом SVM — алгоритме, которому уже несколько десятков лет. И знаете что? Это работает настолько хорошо, что используется даже в академических системах. В этой статье мы разберём, как устроено его решение, почему «примитивный» подход работает лучше сложных нейросетей, и когда вам тоже стоит выбрать SVM вместо трансформера. Давайте разбираться!

    habr.com/ru/articles/990386/

    #SVM #Andrej_Karpathy #TFIDF #машинное_обучение #Support_Vector_Machine #нейросети #алгоритмы_классификации

  33. Почему Andrej Karpathy использует SVM в 2026 году (и вам тоже стоит)

    На arXiv каждый день публикуются сотни статей по машинному обучению. Читать всё — нереально, а пропустить что-то важное — обидно. Andrej Karpathy, бывший Director of AI в Tesla и соавтор курса Stanford CS231n, решил эту проблему неожиданным способом. Он выбрал не BERT, не GPT и не какой-нибудь модный трансформер. Он остановился на добром старом SVM — алгоритме, которому уже несколько десятков лет. И знаете что? Это работает настолько хорошо, что используется даже в академических системах. В этой статье мы разберём, как устроено его решение, почему «примитивный» подход работает лучше сложных нейросетей, и когда вам тоже стоит выбрать SVM вместо трансформера. Давайте разбираться!

    habr.com/ru/articles/990386/

    #SVM #Andrej_Karpathy #TFIDF #машинное_обучение #Support_Vector_Machine #нейросети #алгоритмы_классификации

  34. Почему Andrej Karpathy использует SVM в 2026 году (и вам тоже стоит)

    На arXiv каждый день публикуются сотни статей по машинному обучению. Читать всё — нереально, а пропустить что-то важное — обидно. Andrej Karpathy, бывший Director of AI в Tesla и соавтор курса Stanford CS231n, решил эту проблему неожиданным способом. Он выбрал не BERT, не GPT и не какой-нибудь модный трансформер. Он остановился на добром старом SVM — алгоритме, которому уже несколько десятков лет. И знаете что? Это работает настолько хорошо, что используется даже в академических системах. В этой статье мы разберём, как устроено его решение, почему «примитивный» подход работает лучше сложных нейросетей, и когда вам тоже стоит выбрать SVM вместо трансформера. Давайте разбираться!

    habr.com/ru/articles/990386/

    #SVM #Andrej_Karpathy #TFIDF #машинное_обучение #Support_Vector_Machine #нейросети #алгоритмы_классификации

  35. ICYMI: New Search Engine: Relevance Factors & Ranking Explained! #shorts: A new search engine must prioritize factors beyond simple text matches. Considerations include article freshness, popularity, and source to ensure the most relevant results. Location matters for restaurant searches, ensuring nearby options appear first. #searchengine #relevance #TFIDF #location #proximity youtube.com/shorts/nmWQGIm_hmc

  36. ICYMI: New Search Engine: Relevance Factors & Ranking Explained! #shorts: A new search engine must prioritize factors beyond simple text matches. Considerations include article freshness, popularity, and source to ensure the most relevant results. Location matters for restaurant searches, ensuring nearby options appear first. #searchengine #relevance #TFIDF #location #proximity youtube.com/shorts/nmWQGIm_hmc

  37. ICYMI: New Search Engine: Relevance Factors & Ranking Explained! #shorts: A new search engine must prioritize factors beyond simple text matches. Considerations include article freshness, popularity, and source to ensure the most relevant results. Location matters for restaurant searches, ensuring nearby options appear first. #searchengine #relevance #TFIDF #location #proximity youtube.com/shorts/nmWQGIm_hmc

  38. ICYMI: New Search Engine: Relevance Factors & Ranking Explained! #shorts: A new search engine must prioritize factors beyond simple text matches. Considerations include article freshness, popularity, and source to ensure the most relevant results. Location matters for restaurant searches, ensuring nearby options appear first. #searchengine #relevance #TFIDF #location #proximity youtube.com/shorts/nmWQGIm_hmc

  39. New Search Engine: Relevance Factors & Ranking Explained! #shorts: A new search engine must prioritize factors beyond simple text matches. Considerations include article freshness, popularity, and source to ensure the most relevant results. Location matters for restaurant searches, ensuring nearby options appear first. #searchengine #relevance #TFIDF #location #proximity youtube.com/shorts/nmWQGIm_hmc

  40. New Search Engine: Relevance Factors & Ranking Explained! #shorts: A new search engine must prioritize factors beyond simple text matches. Considerations include article freshness, popularity, and source to ensure the most relevant results. Location matters for restaurant searches, ensuring nearby options appear first. #searchengine #relevance #TFIDF #location #proximity youtube.com/shorts/nmWQGIm_hmc

  41. New Search Engine: Relevance Factors & Ranking Explained! #shorts: A new search engine must prioritize factors beyond simple text matches. Considerations include article freshness, popularity, and source to ensure the most relevant results. Location matters for restaurant searches, ensuring nearby options appear first. #searchengine #relevance #TFIDF #location #proximity youtube.com/shorts/nmWQGIm_hmc

  42. New Search Engine: Relevance Factors & Ranking Explained! #shorts: A new search engine must prioritize factors beyond simple text matches. Considerations include article freshness, popularity, and source to ensure the most relevant results. Location matters for restaurant searches, ensuring nearby options appear first. #searchengine #relevance #TFIDF #location #proximity youtube.com/shorts/nmWQGIm_hmc

  43. 🚀 Tôi đã hoàn thànhстер structure một moteur tìm kiếm độc lập bằng Java! Sử dụng算法 TF-IDF và BM25, hỗ trợ token hóa, xóa từ trống, và ranking văn bản. Hoàn hảo bằng Java 21, không dùng thư viện bên outsourcing. versione opensourcerecipes trong GitHub. Learn rao về thông tin trích xuất và cơ sở dữ liệu!
    #SearchEngine #Java #TFIDF #BM25 #OpenSource #LearningProject #TiemKiem #JavaCor #LapTrinh #NgoQuyet

    reddit.com/r/opensource/commen

  44. Lazy-fedi-question... I have a "working"(?) code example of TF-IDF #tfidf using #scikitlearn and I know the main concepts, but all the tutorials I find are a bit — I don't want to be harsh but —crappy... Can someone point me to some nice open resource on it?

  45. Lazy-fedi-question... I have a "working"(?) code example of TF-IDF #tfidf using #scikitlearn and I know the main concepts, but all the tutorials I find are a bit — I don't want to be harsh but —crappy... Can someone point me to some nice open resource on it?

  46. Lazy-fedi-question... I have a "working"(?) code example of TF-IDF #tfidf using #scikitlearn and I know the main concepts, but all the tutorials I find are a bit — I don't want to be harsh but —crappy... Can someone point me to some nice open resource on it?

  47. Lazy-fedi-question... I have a "working"(?) code example of TF-IDF #tfidf using #scikitlearn and I know the main concepts, but all the tutorials I find are a bit — I don't want to be harsh but —crappy... Can someone point me to some nice open resource on it?

  48. Lazy-fedi-question... I have a "working"(?) code example of TF-IDF #tfidf using #scikitlearn and I know the main concepts, but all the tutorials I find are a bit — I don't want to be harsh but —crappy... Can someone point me to some nice open resource on it?

  49. Recently I've combined various functions which I've been using in other projects (e.g. my personal PKM toolchain) and published them as new library thi.ng/text-analysis for better re-use:

    - customizable, composable & extensible tokenization (transducer based)
    - ngram generation
    - Porter-stemming & stopword removal
    - vocabulary (bi-directional index) creation
    - dense & sparse multi-hot vector encoding/decoding
    - histograms (incl. sorted versions)
    - tf-idf (term frequency & inverse document frequency), multiple strategies
    - k-means clustering (with k-means++ initialization & customizable distance metrics)
    - similarity/distance functions (dense & sparse versions)
    - central terms extraction

    The attached code example (also in the project readme) uses this package to creeate a clustering of all ~210 #ThingUmbrella packages, based on their assigned tags/keywords...

    The library is not intended to be a full-blown NLP solution, but I keep on finding myself running into these functions/concepts quite often, and maybe you'll find them useful too...

    #Text #Analysis #Cluster #KMeans #TFIDF #Ngram #Vector #TypeScript #JavaScript

  50. Recently I've combined various functions which I've been using in other projects (e.g. my personal PKM toolchain) and published them as new library thi.ng/text-analysis for better re-use:

    - customizable, composable & extensible tokenization (transducer based)
    - ngram generation
    - Porter-stemming & stopword removal
    - vocabulary (bi-directional index) creation
    - dense & sparse multi-hot vector encoding/decoding
    - histograms (incl. sorted versions)
    - tf-idf (term frequency & inverse document frequency), multiple strategies
    - k-means clustering (with k-means++ initialization & customizable distance metrics)
    - similarity/distance functions (dense & sparse versions)
    - central terms extraction

    The attached code example (also in the project readme) uses this package to creeate a clustering of all ~210 #ThingUmbrella packages, based on their assigned tags/keywords...

    The library is not intended to be a full-blown NLP solution, but I keep on finding myself running into these functions/concepts quite often, and maybe you'll find them useful too...

    #Text #Analysis #Cluster #KMeans #TFIDF #Ngram #Vector #TypeScript #JavaScript

  51. Recently I've combined various functions which I've been using in other projects (e.g. my personal PKM toolchain) and published them as new library thi.ng/text-analysis for better re-use:

    - customizable, composable & extensible tokenization (transducer based)
    - ngram generation
    - Porter-stemming & stopword removal
    - vocabulary (bi-directional index) creation
    - dense & sparse multi-hot vector encoding/decoding
    - histograms (incl. sorted versions)
    - tf-idf (term frequency & inverse document frequency), multiple strategies
    - k-means clustering (with k-means++ initialization & customizable distance metrics)
    - similarity/distance functions (dense & sparse versions)
    - central terms extraction

    The attached code example (also in the project readme) uses this package to creeate a clustering of all ~210 #ThingUmbrella packages, based on their assigned tags/keywords...

    The library is not intended to be a full-blown NLP solution, but I keep on finding myself running into these functions/concepts quite often, and maybe you'll find them useful too...

    #Text #Analysis #Cluster #KMeans #TFIDF #Ngram #Vector #TypeScript #JavaScript

  52. Recently I've combined various functions which I've been using in other projects (e.g. my personal PKM toolchain) and published them as new library thi.ng/text-analysis for better re-use:

    - customizable, composable & extensible tokenization (transducer based)
    - ngram generation
    - Porter-stemming & stopword removal
    - vocabulary (bi-directional index) creation
    - dense & sparse multi-hot vector encoding/decoding
    - histograms (incl. sorted versions)
    - tf-idf (term frequency & inverse document frequency), multiple strategies
    - k-means clustering (with k-means++ initialization & customizable distance metrics)
    - similarity/distance functions (dense & sparse versions)
    - central terms extraction

    The attached code example (also in the project readme) uses this package to creeate a clustering of all ~210 #ThingUmbrella packages, based on their assigned tags/keywords...

    The library is not intended to be a full-blown NLP solution, but I keep on finding myself running into these functions/concepts quite often, and maybe you'll find them useful too...

    #Text #Analysis #Cluster #KMeans #TFIDF #Ngram #Vector #TypeScript #JavaScript

  53. Recently I've combined various functions which I've been using in other projects (e.g. my personal PKM toolchain) and published them as new library thi.ng/text-analysis for better re-use:

    - customizable, composable & extensible tokenization (transducer based)
    - ngram generation
    - Porter-stemming & stopword removal
    - vocabulary (bi-directional index) creation
    - dense & sparse multi-hot vector encoding/decoding
    - histograms (incl. sorted versions)
    - tf-idf (term frequency & inverse document frequency), multiple strategies
    - k-means clustering (with k-means++ initialization & customizable distance metrics)
    - similarity/distance functions (dense & sparse versions)
    - central terms extraction

    The attached code example (also in the project readme) uses this package to creeate a clustering of all ~210 #ThingUmbrella packages, based on their assigned tags/keywords...

    The library is not intended to be a full-blown NLP solution, but I keep on finding myself running into these functions/concepts quite often, and maybe you'll find them useful too...

    #Text #Analysis #Cluster #KMeans #TFIDF #Ngram #Vector #TypeScript #JavaScript

  54. Okay, Back of the napkin math:
    - There are probably 100 million sites and 1.5 billion pages worth indexing in a #search engine
    - It takes about 1TB to #index 30 million pages.
    - We only care about text on a page.

    I define a page as worth indexing if:
    - It is not a FAANG site
    - It has at least one referrer (no DD Web)
    - It's active

    So, this means we need 40TB of fast data to make a good index for the internet. That's not "runs locally" sized, but it is nonprofit sized.

    My size assumptions are basically as follows:
    - #URL
    - #TFIDF information
    - Text #Embeddings
    - Snippet

    We can store an index for 30kb. So, for 40TB we can store an full internet index. That's about $500 in storage.

    Access time becomes a problem. TFIDF for the whole internet can easily fit in ram. Even with #quantized embeddings, you can only fit 2 million per GB in ram.

    Assuming you had enough RAM it could be fast: TF-IDF to get 100 million candidated, #FAISS to sort those, load snippets dynamically, potentially modify rank by referers etc.

    6 128 MG #Framework #desktops each with 5tb HDs (plus one raspberry pi to sort the final condidates from the six machines) is enough to replace #Google. That's about $15k.

    In two to three years this will be doable on a single machine for around $3k.

    By the end of the decade it should be able to be run as an app on a powerful desktop

    Three years after that it can run on a #laptop.

    Three years after that it can run on a #cellphone.

    By #2040 it's a background process on your cellphone.

  55. Okay, Back of the napkin math:
    - There are probably 100 million sites and 1.5 billion pages worth indexing in a #search engine
    - It takes about 1TB to #index 30 million pages.
    - We only care about text on a page.

    I define a page as worth indexing if:
    - It is not a FAANG site
    - It has at least one referrer (no DD Web)
    - It's active

    So, this means we need 40TB of fast data to make a good index for the internet. That's not "runs locally" sized, but it is nonprofit sized.

    My size assumptions are basically as follows:
    - #URL
    - #TFIDF information
    - Text #Embeddings
    - Snippet

    We can store an index for 30kb. So, for 40TB we can store an full internet index. That's about $500 in storage.

    Access time becomes a problem. TFIDF for the whole internet can easily fit in ram. Even with #quantized embeddings, you can only fit 2 million per GB in ram.

    Assuming you had enough RAM it could be fast: TF-IDF to get 100 million candidated, #FAISS to sort those, load snippets dynamically, potentially modify rank by referers etc.

    6 128 MG #Framework #desktops each with 5tb HDs (plus one raspberry pi to sort the final condidates from the six machines) is enough to replace #Google. That's about $15k.

    In two to three years this will be doable on a single machine for around $3k.

    By the end of the decade it should be able to be run as an app on a powerful desktop

    Three years after that it can run on a #laptop.

    Three years after that it can run on a #cellphone.

    By #2040 it's a background process on your cellphone.

  56. Okay, Back of the napkin math:
    - There are probably 100 million sites and 1.5 billion pages worth indexing in a #search engine
    - It takes about 1TB to #index 30 million pages.
    - We only care about text on a page.

    I define a page as worth indexing if:
    - It is not a FAANG site
    - It has at least one referrer (no DD Web)
    - It's active

    So, this means we need 40TB of fast data to make a good index for the internet. That's not "runs locally" sized, but it is nonprofit sized.

    My size assumptions are basically as follows:
    - #URL
    - #TFIDF information
    - Text #Embeddings
    - Snippet

    We can store an index for 30kb. So, for 40TB we can store an full internet index. That's about $500 in storage.

    Access time becomes a problem. TFIDF for the whole internet can easily fit in ram. Even with #quantized embeddings, you can only fit 2 million per GB in ram.

    Assuming you had enough RAM it could be fast: TF-IDF to get 100 million candidated, #FAISS to sort those, load snippets dynamically, potentially modify rank by referers etc.

    6 128 MG #Framework #desktops each with 5tb HDs (plus one raspberry pi to sort the final condidates from the six machines) is enough to replace #Google. That's about $15k.

    In two to three years this will be doable on a single machine for around $3k.

    By the end of the decade it should be able to be run as an app on a powerful desktop

    Three years after that it can run on a #laptop.

    Three years after that it can run on a #cellphone.

    By #2040 it's a background process on your cellphone.

  57. Okay, Back of the napkin math:
    - There are probably 100 million sites and 1.5 billion pages worth indexing in a #search engine
    - It takes about 1TB to #index 30 million pages.
    - We only care about text on a page.

    I define a page as worth indexing if:
    - It is not a FAANG site
    - It has at least one referrer (no DD Web)
    - It's active

    So, this means we need 40TB of fast data to make a good index for the internet. That's not "runs locally" sized, but it is nonprofit sized.

    My size assumptions are basically as follows:
    - #URL
    - #TFIDF information
    - Text #Embeddings
    - Snippet

    We can store an index for 30kb. So, for 40TB we can store an full internet index. That's about $500 in storage.

    Access time becomes a problem. TFIDF for the whole internet can easily fit in ram. Even with #quantized embeddings, you can only fit 2 million per GB in ram.

    Assuming you had enough RAM it could be fast: TF-IDF to get 100 million candidated, #FAISS to sort those, load snippets dynamically, potentially modify rank by referers etc.

    6 128 MG #Framework #desktops each with 5tb HDs (plus one raspberry pi to sort the final condidates from the six machines) is enough to replace #Google. That's about $15k.

    In two to three years this will be doable on a single machine for around $3k.

    By the end of the decade it should be able to be run as an app on a powerful desktop

    Three years after that it can run on a #laptop.

    Three years after that it can run on a #cellphone.

    By #2040 it's a background process on your cellphone.

  58. Okay, Back of the napkin math:
    - There are probably 100 million sites and 1.5 billion pages worth indexing in a #search engine
    - It takes about 1TB to #index 30 million pages.
    - We only care about text on a page.

    I define a page as worth indexing if:
    - It is not a FAANG site
    - It has at least one referrer (no DD Web)
    - It's active

    So, this means we need 40TB of fast data to make a good index for the internet. That's not "runs locally" sized, but it is nonprofit sized.

    My size assumptions are basically as follows:
    - #URL
    - #TFIDF information
    - Text #Embeddings
    - Snippet

    We can store an index for 30kb. So, for 40TB we can store an full internet index. That's about $500 in storage.

    Access time becomes a problem. TFIDF for the whole internet can easily fit in ram. Even with #quantized embeddings, you can only fit 2 million per GB in ram.

    Assuming you had enough RAM it could be fast: TF-IDF to get 100 million candidated, #FAISS to sort those, load snippets dynamically, potentially modify rank by referers etc.

    6 128 MG #Framework #desktops each with 5tb HDs (plus one raspberry pi to sort the final condidates from the six machines) is enough to replace #Google. That's about $15k.

    In two to three years this will be doable on a single machine for around $3k.

    By the end of the decade it should be able to be run as an app on a powerful desktop

    Three years after that it can run on a #laptop.

    Three years after that it can run on a #cellphone.

    By #2040 it's a background process on your cellphone.

  59. Identified words most associated with a few BlueSky posters. Trained with a tiny dataset of ~2k posts from 7 people via a strongly regularized #ML #NLP #TF-IDF logistic regression model. Last picture shows words for 4 posters: me, space imagery, physician, and original inspiration for idea.

  60. Identified words most associated with a few BlueSky posters. Trained with a tiny dataset of ~2k posts from 7 people via a strongly regularized #ML #NLP #TF-IDF logistic regression model. Last picture shows words for 4 posters: me, space imagery, physician, and original inspiration for idea.