https://www.europesays.com/2980573/ Weng Jiayi, OpenAI Post-Training Engineer, Proposes New Paradigm Hypothesis for Agentic AI #AgenticAI #AgenticArtificialIntelligence #AI #AIImprovement #ArtificialIntelligence #AtariBreakout #Atari57 #ChatGPT #Codex #HeuristicLearning #HeuristicSystem #HumanNormalizedScore #LargeModels #ProgrammaticReinforcementLearning #ReinforcementLearning #WengJiayi
#largemodels — Public Fediverse posts
Live and recent posts from across the Fediverse tagged #largemodels, aggregated by home.social.
-
Benchmark hiệu năng mô hình DeepSeek 671B trên 8 x RTX PRO 6000S sử dụng llama.cpp (layer split mode). Ở định dạng Q4_K_M, tốc độ đạt ~1015 t/s (prefill) và 40.74 t/s (generation). Với Q8_0, tốc độ cao hơn nhưng chiếm nhiều VRAM (~664GB). Hiệu suất thay đổi theo độ dài context (4k–64k). Dữ liệu hỗ trợ lựa chọn cấu hình phù hợp cho LLAMA cục bộ. #DeepSeek #llama.cpp #AI #HPC #DeepSeek671B #MôHìnhLớn #AIInference #DeepSeek #llama.cpp #AI #HighPerformanceComputing #LargeModels #AIInference
https:/
-
Real-time action chunking with large models
https://www.pi.website/research/real_time_chunking
#HackerNews #RealTimeActionChunking #LargeModels #AIResearch #MachineLearning #Innovation
-
Are there any large models doing 3D art that can be 3D printed?