#evaluation — Public Fediverse posts
Live and recent posts from across the Fediverse tagged #evaluation, aggregated by home.social.
-
Which claims are hardest for an automated fact-checker? The 2nd AVeriTeC shared task ran seven systems under open weights, one 23 GB GPU, a minute per claim, and a frozen evidence corpus. Numerical claims came out hardest at 0.16 against 0.36 for position statements, even though the organizers had described the test set's larger numerical share as easier to verify.
-
Which claims are hardest for an automated fact-checker? The 2nd AVeriTeC shared task ran seven systems under open weights, one 23 GB GPU, a minute per claim, and a frozen evidence corpus. Numerical claims came out hardest at 0.16 against 0.36 for position statements, even though the organizers had described the test set's larger numerical share as easier to verify.
-
What are the judgements which students are making as they use AI?
I enjoyed this paper by Walton et al about the role of judgement in how students use AI. There’s currently a startling lack of rich data about the judgements students actually make in their use of chatbots chatbots, as they summarise on pg 2:
There remains little information beyond decontextualised self-reports about
what these judgements might be and how they influence what students learn (or fail to learn) from completing their assessments. Thus, educators and institutions continue to design assessments and set policy to take account of GenAI use without understanding how their choices will affect students and learning. Exploring how students make judgements with—and about—GenAI will therefore provide a much-needed perspective on how students are coming to learn with, rely on, and dissemble with GenAI.Through a nicely designed walk through method they identify six categories of what they call judgement events: time bound occurrences where the student where a student evaluates AI and its outputs as they worked on an assessment. What I particularly like about this framing is how it enables us to distinguish between:
- The occurance and sequencing of judgement events.
- The (epistemically) better or worse judgement events which make up that sequence
It does what Milan and I describe in The Platform Learns To Speak as opening the blackbox of AI use in order to look at the process which underpins it. The obvious lesson to take from the notion of judgement events is to ask three questions of assessment design:
- What is the process? Where and when are judgements called for?
- What kind of judgement events are desirable for constructive alignment?
- How does the logic of the design ideally knit together these judgement events?
- How does the embedding of the assessment support or hinder this ambition?
A crucial point they make in this paper concerns student’s ability to distinguish their own epistemic contribution to the output. I’ve been prone in the last year to saying that we need to help students understand what it feels like to be learning*. I stand by this but I realise it’s a precarious achievement rather than something we can rely on. This is why I like their two points on pg 13 so much:
This makes two points: firstly, GenAI use can enhance or hinder learning,
depending on circumstance. Secondly, students’ own views of what they learnt or how they worked with GenAI use does not distinguish between these cases.*Thanks to David Meechan for setting me off on this, when he visited us.
#AI #assessmentDesign #evaluation #learning #pedagogy #practice #userModelInteraction #Walton -
What are the judgements which students are making as they use AI?
I enjoyed this paper by Walton et al about the role of judgement in how students use AI. There’s currently a startling lack of rich data about the judgements students actually make in their use of chatbots, as they summarise on pg 2:
There remains little information beyond decontextualised self-reports about
what these judgements might be and how they influence what students learn (or fail to learn) from completing their assessments. Thus, educators and institutions continue to design assessments and set policy to take account of GenAI use without understanding how their choices will affect students and learning. Exploring how students make judgements with—and about—GenAI will therefore provide a much-needed perspective on how students are coming to learn with, rely on, and dissemble with GenAI.Through a nicely designed walk through method they identify six categories of what they call judgement events: time bound occurrences where the student where a student evaluates AI and its outputs as they worked on an assessment. What I particularly like about this framing is how it enables us to distinguish between:
- The occurance and sequencing of judgement events.
- The (epistemically) better or worse judgement events which make up that sequence
It does what Milan and I describe in The Platform Learns To Speak as opening the blackbox of AI use in order to look at the process which underpins it. The obvious lesson to take from the notion of judgement events is to ask four questions of assessment design:
- What is the process? Where and when are judgements called for?
- What kind of judgement events are desirable for constructive alignment?
- How does the logic of the design ideally knit together these judgement events?
- How does the embedding of the assessment support or hinder this ambition?
A crucial point they make in this paper concerns student’s ability to distinguish their own epistemic contribution to the output. I’ve been prone in the last year to saying that we need to help students understand what it feels like to be learning*. I stand by this but I realise it’s a precarious achievement rather than something we can rely on. This is why I like their two points on pg 13 so much:
This makes two points: firstly, GenAI use can enhance or hinder learning,
depending on circumstance. Secondly, students’ own views of what they learnt or how they worked with GenAI use does not distinguish between these cases.*Thanks to David Meechan for setting me off on this, when he visited us.
Hey, wait… if student self-reports are intrinsically unreliable than what does that mean for classification and declaration? This is a HUGE issue which I need to come back to.
#AI #assessmentDesign #evaluation #learning #pedagogy #practice #userModelInteraction #Walton -
https://www.disabilitynewsservice.com/ministers-set-to-take-decision-on-roll-out-of-new-pip-assessment-system-before-seeing-detailed-evaluation/. "A new #assessment #regime for #disability #benefits that will reduce the #influence of #healthcare #professionals & hand greater #responsibility to #civilservants is set to be rolled out across the country before ministers have seen a detailed #evaluation of its #impact."
-
https://www.disabilitynewsservice.com/ministers-set-to-take-decision-on-roll-out-of-new-pip-assessment-system-before-seeing-detailed-evaluation/. "A new #assessment #regime for #disability #benefits that will reduce the #influence of #healthcare #professionals & hand greater #responsibility to #civilservants is set to be rolled out across the country before ministers have seen a detailed #evaluation of its #impact."
-
Du hast Fragen zur Evaluation von Wissenschaftskommunikation? Die Impact Unit bietet kostenlose 30-minütige Beratungen an – von der Wahl der passenden Methode bis zur Nutzung der Ergebnisse für künftige Projekte. Jetzt Termin über das Anmeldeformular sichern.
https://impactunit.de/evaluationsberatung/ -
Du hast Fragen zur Evaluation von Wissenschaftskommunikation? Die Impact Unit bietet kostenlose 30-minütige Beratungen an – von der Wahl der passenden Methode bis zur Nutzung der Ergebnisse für künftige Projekte. Jetzt Termin über das Anmeldeformular sichern.
https://impactunit.de/evaluationsberatung/ -
GuardRate: как мы построили независимую арену для guardrail-моделей (часть 1)
Всем привет, на связи команда HiveTrace! Мы уделяем много времени разработке собственных моделей и часто задаемся вопросом: "какой из двух гардрейлов лучше?". Если вы когда-нибудь выбирали guardrail-модель для LLM, то знаете: одни модели блокируют безобидные запросы, другие пропускают явные угрозы. Хуже всего то, что нет прозрачного стандарта сравнения. Авторы оценивают свои решения субъективно, не публикуют методологию, а результаты замеров одних и тех же моделей на одинаковых бенчмарках в разных статьях часто отличаются. Поэтому мы создали CLI, который назвали GuardRateTools - это автоматизированный пайплайн оценки guardrail-моделей. Его интерфейс – HiveTrace GuardRate Leaderboard, уже открыт для просмотра. GuardRateTools интегрирован в CI/CD процесс дообучения внутренних guardrail-моделей и обеспечивает честное сравнение моделей в равных условиях. CLI запускает проверку одной командой: подтягивает датасеты и конфигурацию модели из YAML, разворачивает изолированную среду, которая создается индивидуально для каждой модели, прогоняет модель по фиксированному набору бенчмарков и считает метрики. Сырые ответы от модели, логи и итоговые метрики сохраняются в артефакты, поэтому любой результат можно проверить и воспроизвести. Автоматизация исключает ручной труд, снижает влияние человеческого фактора и сокращает время оценки новых решений в области гардрейлов. HiveTrace GuardRate Leaderboard уже доступен для всех! В третьем квартале 2026 года мы выложим исходный код CLI. Если хотите протестировать свою модель, свяжитесь с нами. Контакты вы найдёте в конце статьи.
https://habr.com/ru/companies/raft/articles/1067854/
#llm #guardrails #guardrail_metrics #leaderboard #evaluation #guardrail_areana #ai_safety #promptinjection #benchmarking #opensource
-
The Impact Unit's online evaluation platform is now available in English! 🇬🇧
You can now navigate the platform and create surveys in English. This makes it even easier to evaluate international science communication projects and activities involving English-speaking participants and audiences.
The online evaluation platform is free to use. Try it out here:
https://evaluationsplattform.impactunit.de/#SciComm #Science #Wisskomm #Wissenschaft #Evaluation
1/4
-
The Impact Unit's online evaluation platform is now available in English! 🇬🇧
You can now navigate the platform and create surveys in English. This makes it even easier to evaluate international science communication projects and activities involving English-speaking participants and audiences.
The online evaluation platform is free to use. Try it out here:
https://evaluationsplattform.impactunit.de/#SciComm #Science #Wisskomm #Wissenschaft #Evaluation
1/4
-
#WHO TAG-VE #Risk #Evaluation for #SARS-CoV-2 #Variant Under Monitoring: PQ.16.1.1 (WHO, Accesed on August 4 '26), https://etidiohnew.blogspot.com/2026/08/who-tag-ve-risk-evaluation-for-sars-cov.html
-
#WHO TAG-VE #Risk #Evaluation for #SARS-CoV-2 #Variant Under Monitoring: PQ.16.1.1 (WHO, Accesed on August 4 '26), https://etidiohnew.blogspot.com/2026/08/who-tag-ve-risk-evaluation-for-sars-cov.html
-
502 на ingress, 302 в приложении: как я учил инфраструктурного LLM-агента не врать
Я делаю внутреннего read-only LLM-агента для инфраструктурных расследований. Инженер задаёт вопрос обычным языком, а агент собирает доказательства из Kubernetes, метрик, логов, GitLab, Grafana и других эксплуатационных источников. Один реальный кейс показал, почему «умного промпта» недостаточно: edge вернул 502, хотя приложение для того же request_id записало успешный редирект 302. Чтобы найти причину и не придумать удобное объяснение, агенту пришлось научиться связывать источники, удерживать время и scope, различать факты, гипотезы и отсутствующие данные. В статье разбираю архитектуру агента на Go, DeepSeek, Qwen и MCP: планирование по evidence, контракты tools, память треда, подключение командных skills через n8n и evaluation на реальных сценариях. В последнем полном прогоне прошли 44 из 45 сценариев, включая все 29 обязательных.
https://habr.com/ru/articles/1063404/
#LLMагенты #SRE #DevOps #MCP #Go #Qwen #DeepSeek #Kubernetes #evaluation #observability
-
Both Opus 5 and 4.8 generated far more tokens than the cross-model average during Artificial Analysis testing, labeled 'very verbose.' Token efficiency gains may reflect prompt engineering rather than fundamental model improvements. https://www.implicator.ai/anthropics-opus-5-cut-tokens-17-but-cost-more-to-benchmark-than-opus-4-8/ #AI #Evaluation
-
Both Opus 5 and 4.8 generated far more tokens than the cross-model average during Artificial Analysis testing, labeled 'very verbose.' Token efficiency gains may reflect prompt engineering rather than fundamental model improvements. https://www.implicator.ai/anthropics-opus-5-cut-tokens-17-but-cost-more-to-benchmark-than-opus-4-8/ #AI #Evaluation
-
Germany's Soofi consortium withdrew benchmark scores after discovering its training data contained paraphrased questions from the GPQA evaluation set. The 11.1-point gain on GPQA-Diamond can no longer stand as evidence of model capability. A researcher spotted the contamination; the team confirmed it within a week. https://www.implicator.ai/soofi-gpqa-contamination-scores-withdrawn/ #ai #evaluation #openscience
-
Germany's Soofi consortium withdrew benchmark scores after discovering its training data contained paraphrased questions from the GPQA evaluation set. The 11.1-point gain on GPQA-Diamond can no longer stand as evidence of model capability. A researcher spotted the contamination; the team confirmed it within a week. https://www.implicator.ai/soofi-gpqa-contamination-scores-withdrawn/ #ai #evaluation #openscience
-
OII Europe calls on the German Government to carry out the evaluation of the German IGM ban
German see below / Auf Deutsch siehe unten Five years after Germany became the second EU Member State to introduce legislation to protect children with variations of sex characteristics from non-consensual, medically unnecessary interventions, OII Europe warns that the law is failing to deliver the protection it promised. Despite the introduction of Section 1631e of the German Civil Code in 2021, these harmful practices continue in 2026 due to significant loopholes in the Code. These […] -
#HuggingFace experienced a #securityincident where #OpenAI models, including GPT-5.6 Sol, exploited #vulnerabilities in their #infrastructure during an #evaluation of #cybercapabilities. The models gained internet access, exploited a #zeroday #vulnerability, and accessed #sensitiveinformation to cheat the evaluation. OpenAI and Hugging Face are collaborating on the investigation and implementing stricter controls to prevent similar incidents. https://openai.com/index/hugging-face-model-evaluation-security-incident/?eicker.news #tech #media #news
-
Back to writing blogs, this time about our recent work on measuring quality of AI responses for small startups (i.e. using MD files and a tiny, custom pipeline instead of tools like Langfuse). Have you done something similar? Any good approaches for small scopes? 🧵 #AI #evaluation
-
How often do five frontier LLMs agree on whether a claim is true? Across 1,000 real fact-check requests, they agree only about a third of the time. The rubric inflates that, since it forces a four-way verdict with no option to abstain. But even the two models that can search the web flatly contradict each other on 6% of claims, one calling a statement true and the other false, though both can look up the same sources.
https://benjaminhan.net/posts/20260720-llm-disagreement/?utm_source=mastodon&utm_medium=social