#llm_evaluation — Public Fediverse posts
Live and recent posts from across the Fediverse tagged #llm_evaluation, aggregated by home.social.
-
7 ошибок в оценке качества LLM‑систем в продакшене
LLM‑фича может месяцами показывать хорошие метрики и при этом регулярно ошибаться на реальных запросах. В статье разберём 7 типичных провалов в оценке качества LLM‑систем: где ломаются эвалы, почему врут дашборды и как выстроить контроль, которому можно доверять в продакшене. Найти ошибки
https://habr.com/ru/companies/otus/articles/1067748/
#LLM #оценка_качества_LLM #LLM_evaluation #LLMasaJudge #наблюдаемость #OpenTelemetry #трассировка #продакшен #метрики_качества #тестирование_LLM
-
Which AI Lies Best? A game theory classic designed by John Nash
https://so-long-sucker.vercel.app/
#ycombinator #AI_deception #AI_benchmark #Gemini_3 #GPT #LLM_evaluation #AI_safety #game_theory #John_Nash #betrayal_game #AI_alignment #machine_learning #artificial_intelligence #AI_behavior #deception_detection -
Which AI Lies Best? LLMs play a 1950s betrayal game by John Nash
https://so-long-sucker.vercel.app/
#ycombinator #AI_deception #AI_benchmark #Gemini_3 #GPT #LLM_evaluation #AI_safety #game_theory #John_Nash #betrayal_game #AI_alignment #machine_learning #artificial_intelligence #AI_behavior #deception_detection