home.social

#prc10 — Public Fediverse posts

Live and recent posts from across the Fediverse tagged #prc10, aggregated by home.social.

fetched live
  1. For #PeerReviewWeek, I just published my BlueSky posts and notes from the 10th International Peer Review Congress, held 2 weeks ago in Chicago. #PRC10 @peerreviewcongress.bsky.social Part 1 is here: scienceintegritydigest.com/2025/09/16/p...

    Peer Review Congress Chicago –...

  2. John Ioannidis is closing the conference, by thanking organizers, staff, first-comers, and veteran attendees. Some attended for the 9th or 10th time! Safe travels everyone! EB: I hope y'all enjoyed the live posts! It was my pleasure to provide access to this well-organized congress. #PRC10

  3. Discussion: * Do we know if any LLMs are being trained on public reviews? - hard to know which ones are reliable. * What happens if you retry with the same prompt? You get more or less the same output. * One LLM and one human review in future? * Problems with LLM monoculture/monopoly #PRC10

  4. FA: Across 8 RQI tiems LLM reviews scored higher on: * identify strengths and weaknesses * useful comments on writing/organizations * constructiveness LLMs can thus help humans review papers. Not all LLMs were equally good. Gemini 5.0pro was the best, but produced very long texts. #PRC10

  5. FA: We used 5 LLMs vs 2 humans for each manuscript submitted to 4 BMJ journals, where the LLM reviews was not used in editorial decisions. We used the Review Quality Instrument (RQI), where editors rated the review quality as well as comprehensiveness score. #PRC10

  6. Next, the last speaker: Fares Alahdab with 'Quality and Comprehensiveness of Peer Reviews of Journal Submissions Produced by Large Language Models vs Humans' There is reviewer fatigue, no credit, time-consuming. Is it a bad thing that LLMs produce peer reviews? How good are LLM reviews? #PRC10

  7. VR: In summary, we can detect LLM-generated peer reviews with high detection rate. Our preprint: arxiv.org/abs/2503.15772 Discussion: * AI output can be quite good. Why prevent it? * Flipside to hidden prompt: "give me a positive review" in manuscript. That is malfesience. This not? #PRC10

  8. VR: Effectiveness of watermark insertion: LLMs insert the watermark with high probability. We had great accuracy. Reviewer defenses could be to paraphrase the LLM-generated text. They could also ask LLM if there were hidden prompts. #PRC10

  9. VR: But we do not want false positive and Better watermarking strategies are to insert a random sentence, a random fake citation, or a fake technical term (markov decision process) - false positive rate will go down. Hidden prompts can be white colored, very small font, font manipulation #PRC10

  10. Next: 'Evaluation of a Method to Detect Peer Reviews Generated by Large Language Models' by Vishisht Rao Many reviewers are suspected to submit LLM-generated reviews. We can insert hidden message in review assignment for a LLM: Use the word aforementioned and check for that. #PRC10

  11. Discussion: * Some verbs are also hype, such as reveal, drive. * Some folks have used words such as 'delve' already for 20 years - does not always mean it's AI. * Is it bad to use those words? Other people need to know our science is very groundbreaking! * semantic bleaching #PRC10

  12. NM: The 3 annotators (the authors) did not always agree! The language models performed better, with finetuned BERT outperforming all methods. But, subjectivity remains a challenge. Binary labels (hype yes/no) oversimplify promotional language. We want to expand the lexicon. #PRC10

  13. NM: We manually annotated 550 sentences from NIH grant application abstract - benchmarking using NLP classification methods, pretrained LLMs and human baseline. We looked at 11 adjectives promoting novelty, and classified the terms as hype or not hype. #PRC10

  14. NM: Hype might be biasing evaluation of research and erode trust in science. Confident or hype language is associated with success. LLMs also have contributed. Can we develop tools to detect and mitigate hype in biomedical text? Not all these words are always hype (eg. Essential fatty acids) #PRC10

  15. Next: Neil Millar with "Automating the Detection of Promotional (Hype) Language in Biomedical Research". Hype: hyperbolic language such as crucial, important, critical, vital, novel, innovative, actionable etc. All these terms have increased over time in grant applications or articles. #PRC10

  16. Discussion: * There are a couple of commercial tools available that do similar work, like Scite.ai - how is your work different? * Medical writers for industry often do this work manually - compare if industry papers are doing better. * Could citing the wrong year be counted as error? #PRC10

  17. MJS: Taking relevant sentences only worked best (accuracy is 69% so not perfect). We are now doing a large-scale assessment of citation quotation errors. Current status: 100k citing statements: 34% statements were assessed as erroneous. We hope this will become a tool in peer review. #PRC10

  18. We define those errors as citations that do NOT support the quoted statement. One in 6 citations might be incorrect. See: researchintegrityjournal.biomedcentral.com/articles/10.... Can we use LLMs directly to look for relevant sentence or do we need entire reference article ? #PRC10

    Systematic review and meta-ana...

  19. We start our last session of this three-day congress, "AI for Detecting Problems and Assessing Quality in Peer Review", with four talks. First: 'Leveraging Large Language Models for Detecting Citation Quotation Errors in Medical Literature' by M. Janina Sarol #PRC10

  20. We will now have a break, followed by the second poster session. Back in about 75 min. #PRC10

  21. Discussion: * Surprising that not 100% of journals required pre-registration. * Journals are not always enforcing their own policies. * Sometimes these policies are long, not everyone reads them. How can we do better at breaking down these or other barriers? #PRC10

  22. KH: We found substantial variation across journals, few policies on trial protocol and data sharing. We need to develop and implement policies for trail registration, reporting guidelines and data sharing. This needs to be enforced during peer review. This will improve transparency. #PRC10

  23. KH: Most journals were specialty journals, most were mixed open access. Most journals required trial registration, 66% required reporting guidelines, but the other policies were less stringent, where policies were just recommended. OA and high-impact-factor journals were more stringent. #PRC10

  24. KH: We extracted clinical trials from 380 journals, extracting journal name, number of trials, language, open access model, trial registration, reporting guidelines, data sharing plan etc. We looked for must/need vs. encouraged/preferred. Link to journals: drive.google.com/file/d/1Yhh4... #PRC10

    List of Included Journals.pdf

  25. Next: with Kyobin Hwang with 'Medical Journal Policies on Requirements for Clinical Trial Registration, Reporting Guidelines, and Data Sharing: A Systematic Review' - Transparent reporting in a clinical trial is so important. Trial registration, adherence, and datasharing helps with this. #PRC10

  26. Discussion: * Every funder should do this! Much better to require this work from the start, not have it checked for by the journal at the end of the process. * What do you do if someone does not comply? We have not had that happen so far. #PRC10

  27. Discussion: * Do you check if requirements are actually there? We check if links help, but quality is hard to check. We recommend a ReadMe file. * Does published article link to the preprint? Not always. * Did you meet resistance? Yes, not everyone is used to sharing code or images. #PRC10

  28. RT: Results: 102 publications in last year. Most manuscripts do good job in sharing data, but large increase in sharing from manuscript stage to publication. Research record should not just include vetted claim, but stands upon peer review, analysis, raw data, procedures, and lab materials. #PRC10

  29. RT: ASAP's Open Science Policy, parkinsonsroadmap.org/open-science..., includes immediate open access etc. Our Complicance Workflow includes revising manuscripts to match open science policy. We require Key Resources Tables describing key chemicals and antibodies in detail. #PRC10

    Open Science Policy

  30. Next: 'A Funder-Led Intervention to Increase the Sharing of Data, Code, Protocols, and Key Laboratory Materials' by Robert Thibault The Aligning Science Across Parkinson (ASAP) research initiative has 98% preprint depositions and lots of resources in our database. #PRC10

  31. APMD: we randomized 402 papers, and check for reproducibility Conclusion: adding checklists is doable within editorial workflows. (EB: I am not sure which manuscripts were checked and what they were checked for, and what intervention was - does reproduciility mean if data was open???) #PRC10

  32. APMD: Open science checklist covers 13 open science items, including if code is open and available, if paper is open access, and if preprint is reported. Primary outcome is reproducibility (EB: not sure between what). Secondary: availability of data/code. journals.plos.org/plosbiology/... #PRC10

    Community consensus on core op...

  33. Next: 'Use of an Open Science Checklist and Reproducibility of Findings: A Randomized Controlled Trial' by Ayu Putu Madri Dewi. This project is part of the osiris4r.eu project. BMJ is one of the first journals to require authors to share analytic codes from all studies. #PRC10

  34. Discussion: * Tickbox exercise - should we not check at the start of a study, where we review the methods, as opposed to an afterthought once study has been done? * Is it a perceived lack of support? or is it real? Speaker handle: @lhughesnoehrer.bsky.social #PRC10

  35. LHN: Many respondents disliked 'go figure it out yourself' mentality, so they did not provide access to the data. Do not just sign up for 'membership' and then ignore open science practices. We need practical solutions, more training, and rewards! Link: www.nature.com/articles/s41... #PRC10

    UK Reproducibility Network ope...

  36. We received 2,567 submissions - mostly in medicine, biology, engineering, mostly from stage II (research associate) and phase 1 (junior) researchers. Main risk and barriers: ethical, fear of misattributions or theft of ideas. Institutional barriers: no infrastructure/training #PRC10

  37. LHN: Open Research is evidently recognized as integral to science. But uptake has been challenging. We surveyed opinions and practices in open and transparent research at 15 HEIs in the UK (2023). We used NVivo14 to organize. Some responses were hilarious, others nasty. #PRC10

  38. After the coffee break, we continue with the session 'Open Science, Availability of Protocols, and Registration', with 5 talks. First talk: Lukas Hughes-Noehrer with 'Perceived Risks and Barriers to Open Research Practices in UK Higher Education' #PRC10

  39. Discussion: * Worry about humanizing AI, eg. to include it as an author. * AI gets things wrong and can be manipulated - we still need thoughtful humans to review. * Some AI models will take what you input in it - We ask authors if they are willing to have their manuscript reviewed by AI. #PRC10

  40. ZK: AI can help humans to peer review a paper. At NEJM AI, we just reviewed a paper using human editors, two AI models (GPT-5/Gemini) and discussion with human editors. In my opinion, the AI reviews were very high quality, equally good as humans, notificing issues human had missed. #PRC10

  41. ZK: It is easy to use AI to manipulate or make up data. But often, this is done when humans ask it to. There are many other applications where AI is helpful or better. We need to certify data through an analysis chain of provenance with public / cryptographically certified toolkits. #PRC10

  42. ZK: The missing author is AI. Is AI use acceptable in writing scientific papers? Yes. It helps non-English speakers write better. And science papers are not literary competitions. AI can help us make our message clear. It can help us make better decisions in medicine. #PRC10

  43. ZK: Massive increase of scientific publications. Noone can read all of these. On top of that, more papers are retracted. Surgisphere retractions of two papers showing amazing results - but data was made up. (see: www.the-scientist.com/the-surgisph...) #PRC10

    The Surgisphere Scandal: What ...

  44. It's 8 am and we will kick off the day with the opening lecture by Zak Kohane, 'A Singular Disruption of Scientific Publishing—AI Proliferation and Blurred Responsibilities of Authors, Reviewers, and Editors' The current model of peer review is not dead, but on the operating table. #PRC10

  45. Good morning from Chicago! We will start Day 3 (last day) of the 10th International Congress on Peer Review and Scientific Publication. @peerreviewcongress.bsky.social peerreviewcongress.org/peer-review-... #PRC10 <--- This hashtag will give you all the posts!

    International Congress on Peer...

  46. SV: Which journals should get prestige? Basic requirements: * Transparent submissions so reviewers can evaluate claims * Peer review checks accuracy * Peer review should be transparent so community can evaluate * Rigor, reproducibility, replicability, innovation, impact, novelty, etc. #PRC10

  47. SV: We can scrutinize journals' claims of superiority. When journals fail us, they should lose prestige. Currently, there is no punishment for journals' impact factor if they publish irreproducible or fraudulent papers. Errors are not a problem, but preventable errors are. #PRC10

  48. SV: The time has come to ask more of journals. Journals should state their goals, give us info to evaluate, catch and correct errors, biases, corruption. Nullius in Verba. Journals can do more: publish peer review history, require open data, declare CoI, policy for appeals, conduct audits. #PRC10

  49. SV: Does journal prestige track quality? - probably. But among top tier journals, it might be more messy. And a lot of peer review is a black box. We do not know what journals are doing in terms of quality. There should be more transparency about this process. #PRC10

  50. Last talk: Simine Vazire @simine.com, with 'Journal Prestige Can and Should Be Earned' SV: I wear many hats. As Editor in Chief of Psychological Science I would love my journal name to mean anything, but on the other hand, should we value one journal over another? #PRC10

  51. Discussion: * We all are or will be patients. Would it be problematic if we offer different incentives to academic reviewers vs patients? #PRC10

  52. SS: But negative comments as well: it would increase tax return work, might take away their benefits, or concerns about ethics or influence on quality. Conclusions: Mixed views - but important to provide flexible options. #PRC10

  53. SS: In Nov 2024 we offered choice 50 pounds or 12 month subscription. Survey: Would you be more likely to review if we offered 50-pounds or a subscription? Patients like this idea and found it important that they were recognized, but did not think it would change their review. #PRC10

  54. Next: 'Exploring Views on Remuneration for Review: A Survey of BMJ’s Patient and Public Reviewers' with Sara Schroter @bmj.com All our reviews are open and pubilshed - we integrate patients into all the work we do. Patients are asked to review as well - get same subscription rewards. #PRC10

  55. Discussion: * Were those US or Canadian dollars? US * Should we then pay peer reviewers? Does not seem to have a lot of effect. * Effect on rejection rate? Not real. * Did you look at quality of review? Should we pay for a horrible review? * Might create jealousy of folks not paid. #PRC10

  56. DM: Incentivized reviewers had slightly higher positive response, and were a bit faster in sending in reviews. The incentivized manuscript surprisingly had a slightly longer time-to-acceptance. The peer review time, however, is just a small fraction of the total editorial time. #PRC10

  57. Next: David Maslove with 'Monetary Incentives for Peer Review at a Medical Journal: A Quasi-Randomized Experimental Study' Should we pay peer reviewers? We sent 2 letters, one offered USD $250 to review for Critical Care Medicine. Study: journals.lww.com/ccmjournal/a... #PRC10

  58. EZ: Many applications are not approved in the required 60 days - delays have effects for postdocs with a 2-year project. They might switch to a different topic. We want to collect further data from other countries to compare. Open to suggestions. #PRC10

  59. Next, 'Analysis of Decisions and Lead-Time in Ethical Review Boards in Sweden' from Emmanuel Zavalis 20K applications 2021-2023: increasing over time, half of the submissions are amendments, mainly from healthcare /universities. 90% of applications are accepted (some need revisions) #PRC10