home.social

#arxiv — Public Fediverse posts

Live and recent posts from across the Fediverse tagged #arxiv, aggregated by home.social.

  1. RAG, поиск и LongContext: почему сложные пайплайны не всегда нужны

    В 2023 году RAG был единственным способом засунуть знания в LLM — контекстное окно было маленьким, часто были галлюцинации. RAG постепенно стал одной из самых популярных технологий, чтобы получить базу знаний, по которой можно искать данные через натуральный язык. Но с другой стороны — RAG-пайплайн тяжёлый и трудозатратный. Компания хочет чат по своей документации. Разработчик говорит: нужен пайплайн — чанкинг, эмбеддинги, векторная база, реранкер. Два-три месяца работы плюс сервис, который надо вечно поддерживать, ради 400-страничной документации. И всё это занимает месяцы, тратит ресурсы ради не такой уж и большой выгоды. Сегодня многие модели держат 1M токенов, а то и больше, а в это окно контекста чаще всего спокойно влезает вся документация. Но если не влезет — то обязательно ли сразу строить весь RAG самому? А как понять, когда есть альтернатива RAG, а когда нет? В этой статье мы разберём, почему RAG стал выбором по умолчанию (и почему это было оправдано), посмотрим, как падает качество при Long Context, сравним подходы и узнаем, что, как и когда использовать.

    habr.com/ru/companies/bothub/a

    #документация #техническая_документация #RAG #LLM #longcontext #AI #ИИ #arxiv #эмбеддинги #поиск

  2. RAG, поиск и LongContext: почему сложные пайплайны не всегда нужны

    В 2023 году RAG был единственным способом засунуть знания в LLM — контекстное окно было маленьким, часто были галлюцинации. RAG постепенно стал одной из самых популярных технологий, чтобы получить базу знаний, по которой можно искать данные через натуральный язык. Но с другой стороны — RAG-пайплайн тяжёлый и трудозатратный. Компания хочет чат по своей документации. Разработчик говорит: нужен пайплайн — чанкинг, эмбеддинги, векторная база, реранкер. Два-три месяца работы плюс сервис, который надо вечно поддерживать, ради 400-страничной документации. И всё это занимает месяцы, тратит ресурсы ради не такой уж и большой выгоды. Сегодня многие модели держат 1M токенов, а то и больше, а в это окно контекста чаще всего спокойно влезает вся документация. Но если не влезет — то обязательно ли сразу строить весь RAG самому? А как понять, когда есть альтернатива RAG, а когда нет? В этой статье мы разберём, почему RAG стал выбором по умолчанию (и почему это было оправдано), посмотрим, как падает качество при Long Context, сравним подходы и узнаем, что, как и когда использовать.

    habr.com/ru/companies/bothub/a

    #документация #техническая_документация #RAG #LLM #longcontext #AI #ИИ #arxiv #эмбеддинги #поиск

  3. Procedural Graphs: Self-Evolving Execution Structures for LLM Agents

    arxiv.org/abs/2609.09153

    #arxiv #llm

  4. Procedural Graphs: Self-Evolving Execution Structures for LLM Agents

    arxiv.org/abs/2609.09153

    #arxiv #llm

  5. Procedural Graphs: Self-Evolving Execution Structures for LLM Agents

    arxiv.org/abs/2609.09153

    #arxiv #llm

  6. Procedural Graphs: Self-Evolving Execution Structures for LLM Agents

    arxiv.org/abs/2609.09153

    #arxiv #llm

  7. Procedural Graphs: Self-Evolving Execution Structures for LLM Agents

    arxiv.org/abs/2609.09153

    #arxiv #llm

  8. #ArXiv Annular liquid jets: exact constraint absorption and finite-time loss of transversality arxiv.org/abs/2609.05457

  9. A “Full House” at the Open Journal of Astrophysics!

    In my recent weekly updates of activity at the Open Journal of Astrophysics, the latest of which is here, I’ve remarked that we had built up quite a backlog of papers accepted but yet to be published. This is no doubt due to August being a holiday for many researchers. Well, we seem to be catching up rather rapidly. So far this week we have published eight new papers, and it’s only Wednesday!

    The recent burst of activity has led to us establishing a first, a week in which we have published a full house of at least one paper in each of the six sections of astro-ph on the arXiv. As a reminder, these are they:

    1. astro-ph.GA – Astrophysics of Galaxies. Phenomena pertaining to galaxies or the Milky Way. Star clusters, HII regions and planetary nebulae, the interstellar medium, atomic and molecular clouds, dust. Stellar populations. Galactic structure, formation, dynamics. Galactic nuclei, bulges, disks, halo. Active Galactic Nuclei, supermassive black holes, quasars. Gravitational lens systems. The Milky Way and its contents
    2. astro-ph.CO – Cosmology and Nongalactic Astrophysics. Phenomenology of early universe, cosmic microwave background, cosmological parameters, primordial element abundances, extragalactic distance scale, large-scale structure of the universe. Groups, superclusters, voids, intergalactic medium. Particle astrophysics: dark energy, dark matter, baryogenesis, leptogenesis, inflationary models, reheating, monopoles, WIMPs, cosmic strings, primordial black holes, cosmological gravitational radiation
    3. astro-ph.EP – Earth and Planetary Astrophysics. Interplanetary medium, planetary physics, planetary astrobiology, extrasolar planets, comets, asteroids, meteorites. Structure and formation of the solar system
    4. astro-ph.HE – High Energy Astrophysical Phenomena. Cosmic ray production, acceleration, propagation, detection. Gamma ray astronomy and bursts, X-rays, charged particles, supernovae and other explosive phenomena, stellar remnants and accretion systems, jets, microquasars, neutron stars, pulsars, black holes
    5. astro-ph.IM – Instrumentation and Methods for Astrophysics. Detector and telescope design, experiment proposals. Laboratory Astrophysics. Methods for data analysis, statistical methods. Software, database design
    6. astro-ph.SR – Solar and Stellar Astrophysics. White dwarfs, brown dwarfs, cataclysmic variables. Star formation and protostellar systems, stellar astrobiology, binary and multiple systems of stars, stellar evolution and structure, coronas. Central stars of planetary nebulae. Helioseismology, solar neutrinos, production and detection of gravitational radiation from stellar systems.

    As of yesterday we had published in five of the six categories and were only missing a paper (in Number 4 – High-Energy Astrophysical Phenomena). Today, right on cue, from the queue, came this paper to complete the set:

    There’s still a couple of days to go and we have more papers in the pipeline so it looks like I’ll have a busy time on Saturday morning when I do the next update!

    #arXiv #DiamondOpenAccess #HighEnergyAstrophysicalPhenomena #OpenJournalOfAstrophysics #TheOpenJournalOfAstrophysics
  10. Rxiv mechanobio
    @Rxiv_mechanobio • Joined: Nov 05, 2024

    I toot papers from #PubMed and preprints from #bioRxiv & #arXiv that seem to pertain to mechanobiology.
    My filter can be tested there on PubMed: link.infini.fr/mechanobiology_
    Improvement suggestions in this filter welcome, both to add or remove spurious results: contact @jocelyn_etienne

    #bot #BotsOfMastodon

  11. How many times I'll need to explain to our dear #Reviewer2 that submitting my own #arxiv paper to a conference is not #selfplagiarism 🤔🤷

    There is a lot of discussion on the use of #AI for #reviewing but, for sure, I'd use #AI to automatically #flag bad reviews (according to a number of criteria, including the above but also based on the lenght of the review, use of toxic words, or even asking for validation on a new ideas paper 😁, another one of my pet peeves)

  12. News from arXiv…

    I saw recently an announcement that arXiv – which recently became a standalone nonprofit organization – now hosts more than three million articles. That message took me to a post on the arXiv blog from which I learnt that said blog is on WordPress which makes it possible to include embedded links from this WordPress blog, i.e.:

    https://blog.arxiv.org/2026/07/09/arxiv-now-hosts-over-3-million-articles/?page_id=1854

    This in turn reminded me to answer a question that people sometimes ask about the Open Journal of Astrophysics, namely what would happen if arXiv went offline? Since OJAp is an arXiv-overlay journal this would appear to be fatal. We have been keeping duplicates of all the papers we publish, and associated metadata, in a local repository so that if arXiv went down we could reistate them fairly quickly. All we would need to do is alter the overlay to point to a different web address: there is no reason at all why the overlay idea can’t be adapted to other forms of repository.

    Recently, however, arXiv decided to do a similar thing for all its articles. Here is a blog post about it:

    https://blog.arxiv.org/2026/02/03/arxiv-future-proofs-access-to-research-with-third-party-digital-preservation/

    I blogged about the first stage of this here.

    Anyway, if you use arXiv at all – which will include most astrophysicists – then it would be a good idea to follow the arXiv blog as there is very important news on it, including information about alterations to the schedule of announcements.

    #arXiv #arXivBlog #arXivMirrors #digitalPreservation #OpenAccess #OpenJournalOfAstrophysics #TheOpenJournalOfAstrophysics
  13. Update. You can now see the preprint itself, on #arXiv.
    arxiv.org/abs/2604.24576

    From the abstract: "We assemble BuyTheBy, a large, annotated dataset of timestamped, text-based paper mill advertisements from seven businesses operating out of seven different countries. The dataset consists of 18,710 individual advertisements, of which 15,839 have prices listed. Among these there are 20,598 positions listed as for sale on 5,567 unique products in 14 different product categories with 51,812 timestamped price data points."

    #Misconduct #PaperMills #ScholComm

  14. "Thousands of shady ads sell paper authorship for cash, large-scale investigation finds."
    science.org/content/article/th

    PS: This article from _Science_ reports important news about paper mills. But apart from that, note that it draws from an #arXiv preprint that hasn't been released yet. It's a preprint preprint, and from _Science_. A nice example of the ongoing #ScholComm transformation.

    #Misconduct #PaperMills

  15. @jonmsterling

    They explain that there are simply too many ultra-low quality reviews posted.

    #bioRxiv has not been accepting to archive reviews for years already (at least since 2017). Their FAQ says it's irrelevant as #preprint because they're not novel in themselves, but I believe part of the reason is also that there's a flooding of dodgy reviews.

    #arXiv

    nature.com/articles/d41586-025

    web.archive.org/web/2017041904

  16. Hello Fediverse! I’m working on quantum gravity/EFT bridging and GW constraints.
    Who are must-follow physics folks here?
    Looking for #QuantumGravity #Gravitation #EFT #ArXiv #OpenScience
    New paper/preprint: (DOI/Zenodo)

  17. @sfmatheson @steveroyle @neuralreckoning @biorxivpreprint

    #biorxiv is very strict about experimental only. I have never been able to get a paper without a methods section past biorxiv.

    I want a preprint server that people will look at. So just putting it up on a website isn't good enough for me. I think of OSF and zenodo as places for data, not manuscripts.

    Still have no freakin' clue what #arxiv is smokin' here.

    Open to ideas.

  18. So if preprint servers reject papers editorially without review or possibility of appeal, how are they different from gatekeeping journals, except without peer review or editorial responsibility.

    (The editors at #arxiv and #biorxiv are unclear, unknown, and apparently answer to no one?)

    #AcademicPublishing #preprints

  19. 🚀 So, apparently, the secret to predicting the future is an equation and a paper number. 📄 Meanwhile, #arXiv is like, "Forget time travel, we just need a #DevOps engineer!" 🤖🔧 Because clearly, the only thing more unpredictable than the future is a job in #open #science. 🌟
    arxiv.org/abs/2505.17989 #predictingthefuture #technology #HackerNews #ngated

  20. Qeios and the Nature of a Journal

    Last week I encountered, for the first time, a website called Qeios.com. This is a platform that does peer review of preprints and then posts those approved with Open Access. It also issues a DOI for approved articles. Qeios is also a member of Crossref so presumably the metadata for these articles is deposited there too.

    You might think this is the same as what the Open Journal of Astrophysics does, but it is a bit different. For one thing, it is not an arXiv overlay journal so the preprints actually appear on the Qeios platform, though I suppose there’s nothing to stop authors posting on arXiv either before or after Qeios. Since most astrophysicists find their research on arXiv, the overlay concept seems more efficient the Qeios one.

    Anyway, my attention was drawn to Qeios by an astrophysicist who had been asked to review an article for Qeios that is already under consideration by OJAp. In our For Authors page there is this:

    No paper should be submitted to The Open Journal of Astrophysics that is already published elsewhere or is being considered for publication by another journal.

    This rule is adopted by many journals and has in the past led to authors being banned for breaking it. Apart from anything else it means that the community is not bombarded with multiple review requests for the same paper (as in the case above). There is an issue of research misconduct, the definition of which varies from one institution to another. For reference here is what it says in Maynooth University’s Research Integrity Policy statement:

    Publication of multiplier papers based on the same set(s) or sub-set(s) of data is not acceptable, except where there is full cross-referencing within the papers. An author who submits substantially similar work to more than one publisher must disclose this to the publishers at the time of submission.

    The document also specifically refers to “artificially proliferating publications” as an example of research misconduct. Authors whose papers do end up in multiple journals could thus find themselves in very hot water with their employers as a consequence.

    Getting back to the specifics of Qeios and OJAp, however, there two questions about whether this rule applicable in this situation. One is that the preprint may have been submitted to Qeios after submission to OJAp, which means the rule as written is not violated. The other is whether Qeios counts as a “another journal” in the first place.

    Instead of going into the definition of what a journal is, I’ll refer you to an old post of mine in which I wrote this:

    I’d say that, at least in my discipline, traditional journals are simply no longer necessary for communicating scientific research. I find all the  papers I need to do my research on the arXiv and most of my colleagues do the same. We simply don’t need old-fashioned journals anymore.  Yet we keep paying for them. It’s time for those of us who believe that  we should spend as much of our funding as we can on research instead of throwing it away on expensive and outdated methods of publication to put an end to this absurd system. We academics need to get the academic publishing industry off our backs.

    The point that I have made many times is that the only thing that journals do of any importance is to organize peer-review. The publishing side of the business is simply unnecessary. Journals do not add value to an article, they just add cost. The one thing they do – peer review – is not done by them but by members of the academic community.

    There is a thread on Bluesky by Ethan Vishniac (Editor-in-Chief of the Astrophysical Journal) about Qeios. There are six parts so please bear with me if I include them all to show context:

    https://bsky.app/profile/ethan-vishniac.bsky.social/post/3lco4s6bfq223

    https://bsky.app/profile/ethan-vishniac.bsky.social/post/3lco4s6bllk23

    https://bsky.app/profile/ethan-vishniac.bsky.social/post/3lco4s6bmks23

    https://bsky.app/profile/ethan-vishniac.bsky.social/post/3lco4s6xjqc23

    https://bsky.app/profile/ethan-vishniac.bsky.social/post/3lco4s6xkpk23

    https://bsky.app/profile/ethan-vishniac.bsky.social/post/3lco4s6xlos23

    This thread repeats much of what I’ve said already, but I’d like to draw your attention to the 4th of these messages, which contains

    Qeios.com takes the position that they are not a journal, but a website that vets papers through peer review. The AAS journals (and as far as I know, all other professional journals) does not regard this as a meaningful distinction.

    I’m not sure what a journal actually is, as I think it is an outmoded concept, but I agree with Ethan Vishniac that to all intents and purposes Qeios is a journal. On the other hand, this quote seems to me to contain a tacit acceptance that the only thing that defines a journal is that it vets papers by peer review, which is the point I made above.

    #academicPublishing #arXiv #OpenJournalOfAstrophysics #Qeios

  21. CW: research review

    S. Singh et al., "XCRYPT: Accelerating Lattice Based Cryptography with Memristor Crossbar Arrays"¹

    This paper makes a case for accelerating lattice-based post quantum cryptography (PQC) with memristor based crossbars, and shows that these inherently error-tolerant algorithms are a good fit for noisy analog MAC operations in crossbars. We compare different NIST round-3 lattice-based candidates for PQC, and identify that SABER is not only a front-runner when executing on traditional systems, but it is also amenable to acceleration with crossbars. SABER is a module-LWR based approach, which performs modular polynomial multiplications with rounding. We map the polynomial multiplications in SABER on crossbars and show that analog dot-products can yield a 1.7−32.5× performance and energy efficiency improvement, compared to recent hardware proposals. This initial design combines the innovations in multiple state-of-the-art works -- the algorithm in SABER and the memristive acceleration principles proposed in ISAAC (for deep neural network acceleration). We then identify the bottlenecks in this initial design and introduce several additional techniques to improve its efficiency. These techniques are synergistic and especially benefit from SABER's power-of-two modulo operation. First, we show that some of the software techniques used in SABER, that are effective on CPU platforms, are unhelpful in crossbar-based accelerators. Relying on simpler algorithms further improves our efficiencies by 1.3−3.6×. Second, we exploit the nature of SABER's computations to stagger the operations in crossbars and share a few variable precision ADCs, resulting in up to 1.8× higher efficiency. Third, to further reduce ADC pressure, we propose a simple analog Shift-and-Add technique, which results in a 1.3−6.3× increase in the efficiency. Overall, our designs achieve 3−15× higher efficiency over initial design, and 3−51× higher than prior work.

    #arXiv #ResearchPapers #CryptographyAcceleration #PostQuantumCryptography #LatticeBasedPQC #Memristors

    __
    ¹ arxiv.org/abs/2302.00095