home.social

#parquet — Public Fediverse posts

Live and recent posts from across the Fediverse tagged #parquet, aggregated by home.social.

  1. When does #Iceberg beat #Parquet+projection on #AWSGlue, and when doesn't ?

    An end-to-end #ETL PoC on #AWS to find out: producer, #Kinesis, two #Firehose paths, two #Glue jobs, #Athena.

    🔮 Spoiler: how the data is read is the key to the choice.

    In the article: every choice with its why, plus a few gems from some Glue experience 😄

    alessandra.bilardi.net/diary/a

    #DiaryOfALazyDeveloper

  2. Munquet 0.2.1 just landed on Flathub 🚀

    Fixed a small race condition when canceling a conversion — turns out the process could finish right before you clicked “Yes” 😅

    Two lines later… all good.

    flathub.org/en/apps/io.gitlab.

    #Flatpak #GTK4 #OpenSource #Parquet #DataScience #Linux #Python #PyArrow

  3. Munquet is now officially on Flathub 🎉

    A native Linux app to convert datasets into Apache Parquet using PyArrow backend. Perfect for data science workflows, analytics, and anyone needing fast local conversions.

    Get it here: flathub.org/en/apps/io.gitlab.

    @gnome @xfce @kde @GTK @linux @flathub

    #apache #pyarrow #datascience #parquet #csv #OpenSource #Python #GNOME #GTK4 #Adwaita

  4. 🚀 Munquet — Convert, merge, rename & validate tabular data into Parquet, fully offline & batch-ready.

    GitLab: gitlab.com/zulfian1732/munquet

    Featured in: @severo 's Awesome Parquet: github.com/severo/awesome-parq 🙏

    #Parquet #OpenSource #Python #GNOME #GTK4 #Adwaita #PyArrow

  5. 🚀 Sneak peek Munquet!
    Convert, merge, rename, and validate tabular data safely into Parquet. Works offline, with batch processing and progress feedback.

    GitLab repo:

    gitlab.com/zulfian1732/munquet

    Flathub release coming soon!

    #Python #GTK4 #GNOME #PyArrow #Parquet #DataScience #Libadwaita

  6. Released scrapy-contrib-bigexporter 1.0.0 (codeberg.org/ZuInnoTe/scrapy-c) - additional export formats for the webscraping framework Scrapy.

    Migrated parquet export from fastparquet to pyarrow as fastparquet is deprecated (docs.dask.org/en/stable/change)

    Migrated orc export from pyorc to pyarrow to reduce the number of dependencies

    #scrapy #crawling #python #parquet #orc #pyarrow #webcrawling #scraping

  7. FERC to start distributing their largest dataset in the least convenient possible format. The EQR is already ~100GB in uncompressed CSVs. In XBRL that would be maybe a TB? With the XML of each filing having to be parsed individually are we going to have to spin up a whole cluster just to get it converted to Parquet in less than a day?

    jdsupra.com/legalnews/ferc-pro

    #OpenData #FERC #Parquet #XBRL #EnergyMastodon