#transformer — Public Fediverse posts
Live and recent posts from across the Fediverse tagged #transformer, aggregated by home.social.
-
Positional encoding или как нейросеть «читает» текст
Большинство нейросетей работают с числами и читать как человек они не умеют, поэтому инженерами были придуманы способы помочь нейросети уловить не только смысл конкретного слова, но и всего предложения. Для достижения такой цели используется positional encoding, то есть помимо отдельных слов нейросети дают их порядок в тексте. В данной статье приведен сравнительный анализ чисто смысловых векторов, линейного базиса (порядковый номер слова), базиса Фурье (синусы и косинусы), комплексного (экспоненты) и полиномы Чебышева для кодирования двух английских палиндромов.
https://habr.com/ru/articles/1072510/
#нейросети #программирование #эмбеддинги #эмбеддинг #позиционный_контекст #data_science #embedding #transformer #трансформеры #нейросеть
-
heretic: Tools for removing censorship in open weight LLMs
https://github.com/p-e-w/heretic
#transformer #censorship #heretic #llm #ai #+ -
Représentations croisées
Deux personnes se tiennent devant la même fenêtre. L’une regarde un arbre, l’autre un ciel qui commence après l’arbre. Elles diront toutes deux : « nous avons vu la même chose » et ce sera faux, même avec la meilleure foi du monde. Aucun œil n’occupe l’espace d’un autre œil. C’est une évidence si triviale qu’on l’oublie aussitôt qu’on la formule. Voir suppose un lieu, et deux corps ne partagent jamais le même lieu au même instant. Ce qui vaut pour […] -
Топ вопросов с NLP собеседований: архитектуры LLM, инференс и оптимизация
На NLP/LLM собеседованиях все чаще проверяют не только знание трансформеров, но и понимание того, как устроены современные GPT-like модели: почему большинство генеративных LLM используют decoder-only архитектуру, чем LLaMA отличается от ванильного Transformer, зачем нужны RoPE, RMSNorm, SwiGLU и GQA. Почему инференс LLM дорогой и какие оптимизации помогают его ускорять: fused kernels, FlashAttention, KV-cache и PagedAttention, continuous batching, speculative decoding, квантизация и дистилляция. Много схем и картинок, а также полный список вопросов с собесов в конце.
https://habr.com/ru/articles/1068944/
#LLM #архитектура_LLM #инференс_LLM #ускорение_LLM #оптимизация_LLM #Transformer #FlashAttention #KVcache #квантизация_LLM #Mixture_of_Experts
-
Doing transformer research recently reminded me how simple the concept of attention is. The transformer really just allowed us to scale the fuck out of compute and data (something previous seq2seq architectures failed at).
I think Dario was right, we were probably on track to create LLMs from the moment we invented the transistor, or "arguably even earlier when we first learned to control fire."
-
Finished my #Lego #Transformer #Soundwave I got for my birthday. Put my original G1 Soundwave next to it for scale.
-
it's totally normal to update your website to add a page for a #transformer from a few years ago at 1am right? well i just did that for legacy burn out on my #neocities lol
https://princess-viola.neocities.org/toyphotos/transformers/legacy-burn-out/
-
Memorization vs Generalization: что действительно умеет языковая модель (TLM) на 2 160 параметров (v1.1.0)
Продолжая тему Крошечной Языковой Модели на Nodejs мы поймем где проходит граница ее возможностей. На примере Крошечной Языковой Модели (TLM) из 2 160 параметров мы проведём воспроизводимый эксперимент и увидим в цифрах: знакомый шаблон - 99,93%, перестановка токенов - 81,69%, неизвестная структура -?%.
https://habr.com/ru/articles/1065160/
#machinelearning #machine_learning #backpropagation #selfattention #autograd #nodejs #javascript #sft #нейронная_сеть #transformer
-
People that still believe, in light of Fable- and Sol-class models, that we're not on the cusp of recursive self improvement, what would change your mind?
- Google's models already improving scheduling
- OpenAI using Sol to improve own inference stack and gain 20% lower serving costs and 15% token generation efficiency -
Transformer Transformer: A Unified Model for Motion-Conditioned Robot Co-Design
https://transformer-transformer.github.io/
Comments: https://news.ycombinator.com/item?id=49093232
#HackerNews #Transformer #Robot #CoDesign #Motion #AI #Robotics
-
Bathroom exhaust fan in zone 1
TL;DR 12V / 150mm bathroom exhaust fans are difficult to find in EU (in online shops), but I found two. If you're in UK, it's Manrose, or Dalap in EU. Intro It looks like there is something German in my Balkan genes: I can't stand stale air indoors, so I open windows all the time and install exhaust fans wherever possible. So I added another one in kid's bathroom. This one was tricky, because the hole through a wall is in Zone 1 (right above the bathtube). Firstly I wanted to install 230V […] -
[Перевод] Языковая Модель без магии: Крошечная Language Model на чистом Node.js
Мы создаем крошечную языковую модель с нуля на чистом Node.js без использования TensorFlow или PyTorch, реализуя нейроны, автоград, эмбеддинги, механизм самовнимания (self-attention), полносвязную сеть (FFN), обратное распространение ошибки и SFT, одновременно наблюдая за тем, как отдельные веса и целые матрицы изменяются в процессе обучения Это не новая GPT и не прод ML...
https://habr.com/ru/articles/1063406/
#machinelearning #machine_learning #backpropagation #selfattention #autograd #nodejs #javascript #sft #нейронная_сеть #transformer
-
"Attention Is All You Need" is a 2017 research paper in #machineLearning authored by eight scientists and engineers working at #Google. The paper introduced a new #deepLearning architecture known as the #transformer, based on the #attentionMechanism proposed in 2014 by Bahdanau et al. The transformer approach it describes has become the main architecture of a wide variety of artificial intelligence systems, including #largeLanguageModels. At the time.
https://www.youtube.com/watch?v=JR8d4PExrXI -
9 turns of Ethernet pair on a ø2cm plastic bobbin makes a very inefficient but working #transformer .
Signal is distinguishable up to 2MHz, and max voltage is 0.3V. Even though the signal generator outputs 3V without load, it seems to choke on the inductivity and gives only 0.3 at output.I think this might be good for transferring data. But how to receive it?