#redpajama — Public Fediverse posts
Live and recent posts from across the Fediverse tagged #redpajama, aggregated by home.social.
-
The #RedPajama #LLM is so painfully close to being truly #OpenSource. Just a few tweaks needed:
- Dropping CommonCrawl/C4 entirely
- Fixing the Gutenberg crawler to stick to public domain books
- Filtering arXiv to return only CC-By(-SA) papers
huggingface.co/datasets/togeth… -
@lifearchitect.ai (30/Oct/2023)
Together #AI has released version of the #RedPajama dataset, with 30T filtered and deduplicated tokens from 84 Common Crawl (web crawl, the Google version is known as Colossal Clean Crawled Corpus or C4) dumps covering 5 languages. This believed to be the largest public dataset ever released (#LLM) training.
At an estimated 125TB of data for 30T tokens, RedPajama-Data-v2 is:
* 2.3× larger than the dataset used to train GPT-4 across 13T tokens (estimated). -
Releasing 3B and 7B #RedPajama-#INCITE family of models including base, instruction-tuned & chat models — #TOGETHER
"The biggest takeaway is the demonstration that performant #LLMs can be built quickly by the open-source community. This work builds on top of our 1.2 trillion token RedPajama dataset, EleutherAI’s #Pythia training code, #FlashAttention from #Stanford and #Together, the #HELM benchmarks from Stanford #CRFM and generous support from #MILA, #EleutherAI & #LAION for compute time on the #Summit #supercomputer within the INCITE program award 'Scalable Foundation Models for Transferable Generalist AI'. We believe these kind of open collaborations, at larger scales, will be behind the best #AI systems of the future. "
-
#LLaMA-Nachbau: #RedPajama – erste dezentrale Open-Source-KI mit offenem Datensatz | Developer https://www.heise.de/news/LLaMA-Nachbau-RedPajama-erste-dezentrale-Open-Source-KI-mit-offenem-Datensatz-8971752.html #OpenSource #MaschineLearning #ChatGPT #ArtificialIntelligence
-
Positive that opensource LLMs and AI like StableLM and RedPajama are gaining traction. Really important as alternatives to the completely closed and not-transparant solutions from Microsoft, Google and OpenAI.
https://github.com/stability-AI/stableLM/
https://www.together.xyz/blog/redpajama
#AI #LLM #StableLM #RedPajama #Opensource -
@survey I'm excited about Large Language Models and open source. This isn't the best example, but #RedPajama: https://news.ycombinator.com/item?id=35600860
-
NEW #LLaMA Rebuilt From Scratch - FULL #OpenSource
https://www.youtube.com/watch?v=uF86vcwM6Js
A video for everyone who is too lazy to read the announcement themselves (like me lol).
-
Very interesting claims from #RedPajama. It seems they are about to build a competitive LLM from scratch, with everything to train these models fully reproducibly, from open training data. If true, highly relevant for FAIR research on / with LLMs.
"The most capable foundation models today are closed behind commercial APIs, which limits research, customization, and their use with sensitive data. Fully open-source models hold the promise of removing these limitations, if the open community can close the quality gap between open and closed models. Recently, there has been much progress along this front. In many ways, AI is having its Linux moment. Stable Diffusion showed that open-source can not only rival the quality of commercial offerings like DALL-E but can also lead to incredible creativity from broad participation by communities around the world. A similar movement has now begun around large language models with the recent release of semi-open models like LLaMA, Alpaca, Vicuna, and Koala; as well as fully-open models like Pythia, OpenChatKit, Open Assistant and Dolly.
We are launching RedPajama, an effort to produce a reproducible, fully-open, leading language model. RedPajama is a collaboration between Together, Ontocord.ai, ETH DS3Lab, Stanford CRFM, Hazy Research, and MILA Québec AI Institute. RedPajama has three key components:
Pre-training data, which needs to be both high quality and have broad coverage
Base models, which are trained at scale on this data
Instruction tuning data and models, which improve the base model to make it usable and safe
Today, we are releasing the first component, pre-training data."
Source: www.together.xyz/blog/redpajam…