#localllama — Public Fediverse posts
Live and recent posts from across the Fediverse tagged #localllama, aggregated by home.social.
-
RT @SlimTradeyBaby: Google hat Gemma 4 auf 31 Milliarden Parameter begrenzt. Ein Entwickler sagte: „Scheiß darauf." Er duplizierte nicht einfach nur Schichten. Er führte eine Art neuronale Chirurgie durch: Er fügte völlig neue Schichten ein, die so konstruiert waren, dass sie zunächst wie perfekte Geisterkopien wirkten und die Ausgabe des Modells um genau null veränderten. Dies schuf leeren mentalen Raum innerhalb der KI, damit sie völlig neues Wissen (koreanisches Recht und fortgeschrittene MINT-Fächer) absorbieren konnte, ohne das bereits Gelernte zu zerstören. Dann trainierte er es. Die neuen Schichten blieben nicht passiv. Sie „erwachten" und trugen mehr bei als die ursprünglichen Schichten. Während Google auf Sicherheit setzt, baut das Open-Source-Underground im Stillen die Modelle, die Google nicht bauen wird. Das verändert alles. teddit.net/r/LocalLLaMA/comm… Link Von der LocalLLaMA-Community auf Reddit: Ich habe Gemma4-31B auf 44 Milliarden Parameter (88 Schichten) erweitert — da Google... Entdecke diesen Beitrag und mehr aus der LocalLLaMA-Community auf reddit.com
mehr auf Arint.info
#AICommunity #Gemma4 #LocalLLaMA #MachineLearning #NeuralSurgery #OpenSourceAI #arint_info
-
-
-
https://www.europesays.com/ch-fr/183597/ Ce robot à monter soi-même « plane » en temps réel #amateur #automatisation #cannabis #Capteur #Chatbot #COV #échantillonneur #fumée #IA #IAGénérative #InformationsSurDesOrdinateursPortatifs #LLM #LocalLLaMA #matériel #MQ2 #nouvelles #rapport #Reddit #revues #robot #robotique #Science #ScienceAndTechnology #Sciences #SciencesEtTechnologies #Sparky #Suisse #Technologies #Technology #température #test #top_k
-
So with this and smart use of LLM hosting (vLLM / Ollama etc) I can run the likes of lucidRAG (which needs GPUs) for FREE (well energy prices)...
Local LLM FTW!
#localllama -
So with this and smart use of LLM hosting (vLLM / Ollama etc) I can run the likes of lucidRAG (which needs GPUs) for FREE (well energy prices)...
Local LLM FTW!
#localllama -
Google just released "QAT" versions of their Gemma 4 models. QAT stands for "yeah we know you people don't have enough VRAM so we trained the model knowing you'd quantize it down to 4 bits anyway" and apparently that makes a 4-bit QAT-model perform similar to an 8-bit quantized with previous methods.
This is a game-changer for running LLMs locally. As a first try I'm running unsloth's version of the 12b model released yesterday, and _without_ quantizing the KV-cache and with >128000 byte context it's not even filling up my 16GB VRAM. Prompt processing > 2000t/s and inference at >40 t/s.
-
Google just released "QAT" versions of their Gemma 4 models. QAT stands for "yeah we know you people don't have enough VRAM so we trained the model knowing you'd quantize it down to 4 bits anyway" and apparently that makes a 4-bit QAT-model perform similar to an 8-bit quantized with previous methods.
This is a game-changer for running LLMs locally. As a first try I'm running unsloth's version of the 12b model released yesterday, and _without_ quantizing the KV-cache and with >128000 byte context it's not even filling up my 16GB VRAM. Prompt processing > 2000t/s and inference at >40 t/s.
-
Unsloth Qwen3.6: New coding model of interest if you run AI coding agents locally
https://unsloth.ai/docs/models/qwen3.6
#localllama #claudecode #coding #agents #qwen #ai #+ -
Tapes: Hệ thống giám sát minh bạch cho AI cục bộ, cho phép lưu trữ session, tìm kiếm cuộc trò chuyện, và phục hồi trạng thái như Git. Tích hợp Ollama, mã nguồn mở. Thử ngay! 🔧 #AI #Tech #LocalLLaMA #MáyHọc #PhátTriểnMở
https://www.reddit.com/r/LocalLLaMA/comments/1qsggwk/introducing_tapes_local_transparent_agentic/
-
[vLLM Office Hours #42] Cập nhật chi tiết về tính năng CPU Offloading Connector - 29/01/2026. Tính năng đang được phát triển tích cực, nhiều người dùng chưa hiểu rõ hoặc gặp lỗi trước đây. Xem [link] chi tiết, thảo luận tại subreddit.
#vLLM #AI #CPUOffloading #CôngNghệML #LocalLLaMA #DeepLearning #MáyHọc -
Tối ưu Strix Halo 128GB để chạy mô hình AI 93GB với 64k ngữ cảnh. Cài đặt BIOS (RAM toàn bộ cho CPU), GRUB (amdgpu.gttsize=131072), llama-server -c65536/-b2048, lượng hóa q4_0. Hiệu suất chuẩn 7-9 token/giây. #AI #LLM #StrixHalo #LocalLLaMA #MáyHọc
-
Tác giả đã phát triển Monolith – ứng dụng Windows kết hợp LLM, Stable Diffusion, và tạo âm thanh trong một giao diện. Yêu cầu: Windows, GPU CUDA, Python 3.10+. Bản alpha, mã nguồn mở MIT. Đánh giá tính năng, hiệu năng & tìm thử nghiệm trên AMD/Mac. GitHub: [github.com/Svnse/monolith](https://github.com/Svnse/monolith)
#AI #MachineLearning #LocalLLaMA #StableDiffusion #ÂmThanh #PhátTriểnPhầnMềm #CUDA #Windows #MởMã #ỨngDụngBảnĐịa #AIHiệnĐại #GiaoDịchĐaNăng -
Việc lưu trữ (cache) kết quả embedding giúp tăng tốc xử lý 7.6 lần, từ 7 phút 53 giây xuống còn 62 giây. Bài học: Tránh tái tính toán khi không cần thiết để tiết kiệm GPU và năng lượng. #AI #MáyHọc #Caching #OSS #LocalLLaMA
https://www.reddit.com/r/LocalLLaMA/comments/1qpej60/caching_embedding_outputs_made_my_codebase/
-
Người dùng Reddit báo cáo Jan.ai không nhận diện GPU AMD RX 9070 XT. Họ đã thử, kèm ảnh chụp màn hình cho thấy GPU không được phát hiện và hỏi cách khắc phục. Cần hỗ trợ cấu hình hoặc giải pháp tương thích. #JanAI #AMD #GPU #AI #Tech #CôngNghệ #LocalLLaMA #AI #GPU 🚀
https://www.reddit.com/r/LocalLLaMA/comments/1qk27xm/janai_и_rx_9070xt/
-
RTX 6000 Pro (Blackwell) không khởi động trên mainboard MSI Z790-P. Lỗi: LED EZ Debug nhấp nháy đỏ-vàng rồi tắt. Khắc phục: Dùng iGPU vào BIOS, tắt CSM, bật Above 4GB Decoding, tắt ReBAR. Sau cập nhật BIOS, cài đặt bị reset – cần lặp lại thao tác. Mainboard MSI hỗ trợ GPU yếu, iGPU lỗi khi CSM tắt. Cần thay mainboard khác. #RTX6000Pro #Blackwell #MSI #Z790 #PCIE5 #UEFI #CSM #ReBAR #vietnamese #LocalLLaMA #GPU #BIOS #HACK 🖥️🔧
https://www.reddit.com/r/LocalLLaMA/comments/1qc1isk/rtx_6000_pro_bl
-
@LocalLLaMA Trùng hợp 2 CPU Xeon 2690 V3 + RAM 2133MHz, nâng tối đa 256GB để tự lưu trữ mô hình AI cỡ lớn (LLM). Có thể chạy mô hình 235B như Qwen3? Mời chia sẻ kinh nghiệm & đánh giá hiệu năng! #SelfHosting #LLM #AI #TựNhânViễn #Qwen3 #MáyChủAI #ThửTháchCôngNghệ #LocalLLaMA
https://www.reddit.com/r/LocalLLaMA/comments/1py4nuu/self_hosting_llm_on_multi_cpu_sys_ram_combo/
-
Dự án RAG nhỏ với 16GB VRAM: Người dùng muốn tự lưu trữ mô hình AI để xử lý tài liệu Google trên card đồ họa 16GB, đặt câu hỏi về giới hạn và gợi ý mô hình nhỏ.
#RAG #VRAM #AI #HocMay #CongNgheThôngTin #LocalLLaMA #MôHìnhAIhttps://www.reddit.com/r/LocalLLaMA/comments/1pvv279/small_rag_project_with_16_gb_vram/
-
Xiaomi tung phiên bản mới **MiMo-V2-Flash (309B)**, модель với khả năng xử lý ngôn ngữ cực mạnh, tiến gần hơn đến công nghệ AI hàng đầu. Đây là bước tiến đáng chú ý trong lĩnh vực trí tuệ nhân tạo.
#Xiaomi #AI #TríTuệNhânTạo #CôngNghệ #MIMOV2Flash #TechNews #LocalLLaMA
-
200$ nên làm gì với H100s? Azure cung cấp: 🔥 1x H100: 1.46$/h (eastus2) | 🔥 2x H100: 3.10$/h (northcentralus) | 🔥 8x H100: 16.35$/h (westus3). Hỗ trợ yêu cầu, tinh chỉnh mô hình. #Azure #H100 #AI #LocalLLaMA #MáyHọc #TinhChỉnhMôHình
https://www.reddit.com/r/LocalLLaMA/comments/1pqxw37/what_do_i_do_with_200_for_some_h100s/
-
Exo 1.0 cho phép kết nối Macbook qua cáp Thunderbolt để chạy mô hình AI lớn. Không có giới hạn phần cứng, hỗ trợ đa thiết bị (M1, M2) khác nhau. Cổng Thunderbolt an toàn, tự động phát hiện thiết bị. #Exo10 #MLX #Macbook #AI #KếtNốiMáyTính #TechViet #LocalLLaMA #AppleSilicon #M1M2 #KỹThuậtViễnThông #HợpTácMáyTính
https://www.reddit.com/r/LocalLLaMA/comments/1pqfs1l/exo_10_means_you_can_cluster_mac_studios_for/
-
🌟 Giới thiệu Memora – máy chủ bộ nhớ cục bộ dành cho MCP, được hỗ trợ bởi SQLite, không cần đám mây, có tìm kiếm ngữ nghĩa.
✅ 100% cục bộ – Dữ liệu ở trên máy, không gửi ra ngoài.
🔒 Bảo mật tuyệt đối – Không thu thập dữ liệu, không gọi API (nếu không dùng embeddings).
⚡ Tìm kiếm nhanh – Sử dụng FTS5, kết hợp tìm kiếm từ khóa và ngữ nghĩa.
🧠 Tương thích MCP – Hoạt động với Claude Code, Cursor và các client MCP khác.#LocalLLaMA #AIđịaphương #BảoMật #MCP #NguồnMở #AI #TechVietNam
h
-
Ever wonder how tools analyse gigabytes of data, while you’re fighting to squeeze a few KB of code into an Ollama context window?
The trick isn’t bigger models.
It’s not giving the data to the LLM at all.LLM → SQL → DuckDB → CSV-on-disk
Schema + samples only. Zero data leakage. Sub-100ms queries on 500MB+ files.Local. Private. Scales past RAM.
Full C# walkthrough:
https://www.mostlylucid.net/blog/analysing-large-csv-files-with-local-llms -
**LiteLLM có thực sự ổn định?** Người dùng chia sẻ trải nghiệm về những hiện tượng ngừng hoạt động bất thường khi kết nối các mô hình cục bộ qua LiteLLM. Một số vấn đề được phản hồi: có lệnh hoạt động, có lệnh thì mô hình lại không hiển thị. Cộng đồng đang tranh luận đây là lỗi cá nhân hay vấn đề chung của công cụ này so với Haproxy. #LiteLLM #LocalLLM #AI #CôngNghệ #StabilityAI #MôHìnhLậpDung #LocalLLaMA
(Trích từ thảo luận trên Reddit r/LocalLLaMA)
-
Jan cập nhật v0.7.5 với tiện ích trình duyệt Jan Browser MCP, hỗ trợ đính kèm tệp tin và Flatpak dành cho Linux! Người dùng Chrome có thể cài tiện ích từ Chrome Web Store để tăng tính ổn định khi lướt web. Đồng thời, Linux users cuối cùng được hỗ trợ Flatpak. [Kèm video hướng dẫn chi tiết]. Những cải tiến này giúp trải nghiệm chat và tương tác trở nên thuận tiện hơn. Tìm hiểu thêm tại Flathub hoặc GitHub.
#LocalLLaMA #Jan #AI #PhầnMềmMới #Linux #CôngNghệ
#JanAI #BrowserExtension #FileSuppor -
Chào mừng bạn đến với projekt AI "Atlantis" - công nghệ ẩn بعيدًا về tóc, không nài dữ liệu lên đâu! ✨
Đoạn beta miễn phí, tập trung làm 最佳 trên máy, xử lý thông tin concentración.docx quý giá (memos, email, sức khỏe) mà không cần đомо. Features: briefing sáng, ghi nhớ lâu dài, tìm kiếm thông minh. 🚀
Đầu tư vào riêng tư, không clou đâu. Có Vieux hoặc suggestions? Newsfeed đã sẵn sàngsurface!#AI #PrivacyFocused #TechNews #LocalLLaMA #FeedbackWanted #Jarvis #Vietnamese
Tags: #AI, #Priva -
Chào mừng bạn đến với projekt AI "Atlantis" - công nghệ ẩn بعيدًا về tóc, không nài dữ liệu lên đâu! ✨
Đoạn beta miễn phí, tập trung làm 最佳 trên máy, xử lý thông tin concentración.docx quý giá (memos, email, sức khỏe) mà không cần đомо. Features: briefing sáng, ghi nhớ lâu dài, tìm kiếm thông minh. 🚀
Đầu tư vào riêng tư, không clou đâu. Có Vieux hoặc suggestions? Newsfeed đã sẵn sàngsurface!#AI #PrivacyFocused #TechNews #LocalLLaMA #FeedbackWanted #Jarvis #Vietnamese
Tags: #AI, #Priva -
Flux 2 giờ có thể chạy trên card 24GB VRAM rồi! Dễ dàng hơn nhiều so với tưởng tượng. Ai đang dùng card yếu yên tâm lên nhé!
#Flux2 #LLM #AI #LocalLLaMA #VietnameseAI #trí_tuệ_nhân_tạo #học_máy
https://www.reddit.com/r/LocalLLaMA/comments/1p6ht87/flux_2_can_be_run_on_24gb_vram/
-
Mô hình AI nào chạy Local gần giống GPT-5 nhất? Cộng đồng đang tìm kiếm các mô hình có khả năng suy luận, tìm kiếm web và RAG tốt, đồng thời chia sẻ cấu hình phần cứng cần thiết.
#AI #LLM #GPT5 #LocalLLaMA #VietnameseAIhttps://www.reddit.com/r/LocalLLaMA/comments/1p0ghl7/best_local_model_closest_to_gpt5/
-
Ollama không tận dụng được GPU rời trên Fedora 42 dù đã nhận diện card NVIDIA RTX A5000. CPU hoạt động hết công suất nhưng GPU thì không. Cần tìm cách kích hoạt GPU cho Ollama.
#Ollama #LocalLLaMA #GPU #AI #VietnameseAI #LLM #Fedora #NVIDIA #RTX
https://www.reddit.com/r/LocalLLaMA/comments/1oxd2me/why_isnt_ollama_using_my_dgpu/
-
Phát triển dự án mã nguồn mở như Claude-Code cho DevOps, chạy ngoại tuyến, sử dụng mô hình được đào tạo cục bộ, xử lý các nhiệm vụ như đọc/ghi file, chạy lệnh shell. #DevOps #MãNguồnMở #ClaudeCode #Ollama #LocalLLaMA #DựÁnMới #BetaTester #OfflineAssistant #DevOpsTool
https://www.reddit.com/r/LocalLLaMA/comments/1ox12cj/opensource_local_claudecode_alternative_for/
-
Mô hình DeepSeek-OCR GGUF giờ chạy tốt ngay trên máy local, đơn giản và nhanh chóng! Dễ dàng sử dụng với Hugging Face. Quá trình thiết lập nhanh chóng cho các bạn tự triển khai LLM tại nhà.
#LocalLLaMA #DeepSeekOCR #GGUF #LLM #AI #VietnameseAI #trituexuatzao
-
Playing with compar:IA, the new LLM chatbot comparison arena created by the French government: https://comparia.beta.gouv.fr
I particularly like the "frugal" mode that allows you to compare two randomly chosen small/cheap models against each other! It's great for testing many different models suitable for local hosting on a particular task.
After each chat session, it also shows you an estimate of the energy consumed by each model!
-
Chủ đề: Triển khai mô hình LLM tự huấn luyện thành API trả phí.
Một dev đang tìm cách biến mô hình fine-tune (code review, coding style) chạy trên vLLM thành API để kiếm thêm thu nhập. Các yêu cầu: dashboard, quản lý API key, giới hạn tốc độ, đo lường usage (theo token) và tích hợp thanh toán với Stripe. Có công cụ nào đơn giản hơn Kong không?
#LocalLLaMA #LLM #API #vLLM #AI #SaaS #Stripe #dev #VietnameseAI #HọcMáy
https://www.reddit.com/r/LocalLLaMA/comments/1opjt8f/whats_the_stack_for_going
-
Một người dùng đã thử nghiệm mô hình AI MiniMax-M2 bằng cách yêu cầu nó tạo game Asteroid bằng HTML. Kết quả gây bất ngờ: tốc độ xử lý nhanh, tự động sửa lỗi 100% ở lần thứ hai và tích hợp âm thanh, hiệu ứng hình ảnh mà không cần nhắc. Người dùng đánh giá cao khả năng "suy nghĩ" của mô hình này.
#MiniMaxM2 #AI #Gaming #LocalLLaMA #MôHìnhAI #Gamehttps://www.reddit.com/r/LocalLLaMA/comments/1on8zl6/minimaxm2_asteroid_game_unsloth/
-
Một người dùng đã thử nghiệm mô hình AI MiniMax-M2 bằng cách yêu cầu nó tạo game Asteroid bằng HTML. Kết quả gây bất ngờ: tốc độ xử lý nhanh, tự động sửa lỗi 100% ở lần thứ hai và tích hợp âm thanh, hiệu ứng hình ảnh mà không cần nhắc. Người dùng đánh giá cao khả năng "suy nghĩ" của mô hình này.
#MiniMaxM2 #AI #Gaming #LocalLLaMA #MôHìnhAI #Gamehttps://www.reddit.com/r/LocalLLaMA/comments/1on8zl6/minimaxm2_asteroid_game_unsloth/
-
Để mở rộng quy mô trên 2 DGX Sparks trong một cluster, bạn có thể tận dụng dual 100Gbit QSFP28 links và cấu hình ROCE v2 với layer 3 links. Điều chỉnh NCCL variables để sử dụng ROCE v2 và kết nối cả hai cổng CX7 với switch để tăng băng thông.
#LocalLLaMA #NVIDIA #DGX #clusters #AI #vietnamhttps://www.reddit.com/r/LocalLLaMA/comments/1oieip0/theoretically_scaling_beyond_2_dgx_sparks_in_a/
-
MiniMax-M2 đã có mặt trên Hugging Face nhờ Unsloth! Bản guff có thể sớm ra mắt, mở ra khả năng sử dụng local dễ dàng hơn. 🚀
#LocalLLaMA #AI #MachineLearning #Unsloth #MiniMaxM2 #trituenhantao #hocmayhttps://www.reddit.com/r/LocalLLaMA/comments/1oibaz2/waiting_for_an_unsloth_guff_for_minimaxm2/
-
Do you have any coder LLM recommendations that can run locally with 16 GBs of RAM, other than Qwen-2.5-Coder:7b?
-
Which VSCode extension you'd prefer to try local LLMs or through Openrouter etc. ?