home.social

#smalldata — Public Fediverse posts

Live and recent posts from across the Fediverse tagged #smalldata, aggregated by home.social.

fetched live
  1. 🍗 Oh, the endless saga of "Small Data": where a 2012 MacBook Pro is dusted off to solve the mysteries of the universe! 🤔 Apparently, we've been too busy dreaming of Big Data, only to find out our datasets are about as exciting as watching grass grow in real-time. 🌱✨
    duckdb.org/2025/05/19/the-lost #SmallData #MacBookPro #BigData #MysteriesOfTheUniverse #DataAnalysis #HackerNews #ngated

  2. Big Data, Hochleistungsrechner, KI – diese Themen erhalten derzeit viel Aufmerksamkeit. Aber wie gehen Wissenschaftler*innen Probleme an, wenn die verfügbaren Datenmengen gering sind? Im Sonderforschungsbereich (SFB) „Small Data“ entwickeln Forscher*innen Methoden für solche Small-Data-Anwendungen.
    Im Interview sprechen die Doktorand*innen Maren Hackenberg, Lennart Purucker und Esma Secen über ihre Promotionsprojekte im Bereich Small Data, ihre Erfahrungen mit interdisziplinärer Zusammenarbeit und die erhoffte Wirkung des Sonderforschungsbereichs.
    Zum Interview: uni-freiburg.link/small-data

    #unifreiburg #smalldata #ki #ai #sfb #promotion

  3. Just caught up with the recent Delta Lake webinar,

    > Revolutionizing Delta Lake workflows on AWS Lambda with Polars, DuckDB, Daft & Rust

    Some interesting hints there regarding lightweight processing of big-ish data. Easy to relate to any other framework instead of Lambda, e.g. #ApacheAirflow tasks

    youtu.be/BR9oFD0QMAs

    #dataengineering #datascience #duckdb #daft #polars #pandas #python #spark #deltalake #databricks #airflow #bigdata #smalldata

  4. #SmallData Symposium 2024 am 7.&8.10. zeigt, wie KI auch mit wenig Daten möglich ist.

    Für Pressevertreter*innen besteht am 7.10. um 12:30 Uhr die Möglichkeit, an einem Pressegespräch mit den Vortragenden des ersten Blocks teilzunehmen, um ein tieferes Verständnis der behandelten Themen zu ermöglichen.

    uni-freiburg.de/smalldata-symp

  5. medium.com/ai-in-plain-english

    could be used as a curated data set to tune more precise and specialized models. It is open the door to a and

    is not new, but in a set is relatively new.

    Your ideas, journals, and emotions is are key to the explainability of your hard data points

  6. "While suggesting that there is individual work to be done often gets derided as ‘neoliberal’, it’s often the personal connections we can make and the way we can use data to show people how climate crisis affects their lives that can help us mobilise for broader forms of advocacy at structural levels." (@roopikarisam)

    #ClimateDiary #SmallData

  7. Finally had the time to rent this book (yeah for the orders open to all at the University of Luxembourg!).

    That's what I like about research, the space to explore different ways of communicating and discussing results.

    #SmallData #Academia

  8. "Institutional capacity, useability of data, and evidence-based decision-making for sustainable development in Small Island Developing States"
    A new working paper with Kalim Shah to input into the #SIDS4 meeting via ODI.
    odi.org/en/publications/data-a
     
    #IslandStudies #SmallIslands #BigData #SmallData #SustainableDevelopment #Islands

  9. On the occasion of today's #release, here's a late #introduction ! This is the Mastodon account of the Mégra :megra_wavetable: language, a LISP-y DSL developed by @parkellipsen to make sound and music. It's built around Markov chains and notions of #smalldata . Also, it's a standalone program (implemented in Rust) that comes with its own editor and synthesizer/sampler.

    I took this #patternuary as an opportunity to showcase some of its central features: pixelfed.de/c/6495503601336212

    Version 0.0.12 has been released today!

    If anyone wants to try:
    Docs: megra-doc.readthedocs.io/en/la
    Code: github.com/the-drunk-coder/meg
    Latest release: github.com/the-drunk-coder/meg

    I'd be happy to receive some feedback :)

  10. youtu.be/0T_imbxX5Zg?si=-uXBwj

    and bring an enterprise tech closer to human and brain-friendly system design

  11. - is not about is more about and personal . Curated sets of ideas, thoughts, personal data that focus on values and are value itself. Human brain-friendly data that are in a . will help you create a between and the ocean of information and big data in a digestible and human-friendly way. We desperately need the next generation of

  12. This weekend, I came up with the thought that we need another 2 hashtags: #MessyWeekends and #MedianData. Here's why:

    The inspiration for #MessyWeekends came from seeing some folks regularly having fun with #TidyTuesday. However, since my Tuesdays usually are pretty tight, I will have to find some time on the weekends to play with #Data, #Python, #pandas, #PyData, #DataScience, etc. With a 4.75YO at home, weekends tend to be messy anyway. So that's that.

    As for #MedianData, it is literally between #SmallData and #BigData, and I think a lot can be learned by pouring over it and trying out different kinds of #DataTools.

    What do you think? 😎

  13. Super excited to be at #biohack23 looking forward to connecting #smalldata with big combining chemistry and biology. In any case it's going to be #sparqling.

  14. As co-chair, I'm excited about this 💥 all the individual livestream talks from Citus Con: An Event for Postgres 2023 have published on YouTube 📺

    Including some phenomenal talks 💡 by @simon & @cyberdemn & @tmunro - and others not yet on Mastodon

    Enjoy the playlist of 37 talks & pls tell your friends

    aka.ms/cituscon-playlist

    #CitusCon #PostgreSQL #Citus #SmallData #OpenSource #Patroni #YouTube #DistributedPostgreSQL #HammerDB

  15. mastodon.ie/@Castlebridge/1101 Sometimes it needs to be all about the #SmallData if you want to avoid #IQTrainwreck*

    *#IQTrainWreck is a #DataQuality or #InformationQuality problem that causes an embarrassing but avoidable headline or results in significant impacts on people or organisations. I coined the term back in 2005. See iqtrainwrecks.info for some classic examples or check the hashtag over on the bird site as well.

  16. In my personal experience, smart use of smaller data sets yielded most benefits (kind like 20-80 rule). Often, the cost was in tackling the large information space embedded in small data sets via better features. #bigdata #smalldata motherduck.com/blog/big-data-i

  17. #introduction
    I am an open data access (#ODBC, #JDBC, and #HTTP), integration (#DataVirtualization), and data management (#DBMS and/or #KnowledgeGraph) technologist, enthusiast, and entrepreneur.

    I am passionate about open standards for #GenAI, #AI, #Identity #Authenticity, #SocialMedia (#ActivityPub & #ActivityStreams), and data de-silo-fication initiatives (e.g. #SemanticWeb, #LODCloud, #LinkedData, #RWW, and #SmallData).

    I don't like any kind of silo!

  18. For #SmallData projects (where small is honestly pretty big these days!) I can't recommend this pattern enough: dump EVERYTHING into a #SQLite file - raw data, intermediate representations, final product. This way you get all the advantages of a flat file: transportable, backupable, versionable.

    Then, if/when it comes time to publish the data, all you need to do is copy the file, drop tables you don't need, and ship it!