home.social

#data-engineering — Public Fediverse posts

Live and recent posts from across the Fediverse tagged #data-engineering, aggregated by home.social.

fetched live
  1. Radim Marek's twelve-month AlloyDB evaluation is the most careful public writeup of it so far, and it is not a hit piece.

    His summary: "AlloyDB's compatibility claim holds at the wire protocol, but the underlying engine diverges immediately at the storage layer." In managed AlloyDB the 8KB page is no longer the durable representation of your data; the WAL is, and pages are derived from it.

    boringsql.com/posts/google-all

    pgexperts.com

  2. Two habits keep our AI data pipeline bills sane: batch 20 to 30 items per LLM call instead of one-at-a-time, and use change detection so nothing gets re-enriched unless the source actually changed. go.upgradejs.com/e7w #AI #DataEngineering #LLM

  3. Publications sur les Data Engineer Junior / ingénierie de données pour les juniors : 20k+ vues sur LinkedIn en quelques jours... Faisons du réseau ! Pas de conseils à vendre, juste échanger.

    #DataEngineer #DataEngineering #Junior #Recrutement #RH #Senior #OpenToWork #LinkedIn !

  4. New on the Psycopg blog: row-by-row streaming with server-side cursors 🐘🐍

    One of the most common mistakes with large PostgreSQL datasets is calling fetchall() and loading everything into memory. Server-side cursors solve this cleanly: stream results row by row, control batch size, and use it with async too.

    We walk through how it works, a practical export example, and what to watch out for. Link in replies 👇

  5. New partner onboarding used to eat days of dev time. We rebuilt it as an AI pipeline where the LLM drafts the mapping and deterministic code validates every output before it ships. Same job, minutes of compute instead of days of debugging. Proposal engine, not decision maker.

    go.upgradejs.com/uey

    #AI #DataPipelines #LLM #DataEngineering

  6. Next Wednesday: the engineer who built ColdFront live on the architecture.

    Most #Postgres databases pay SSD prices for data nobody queries. ColdFront gives you #PostgreSQL to #ApacheIceberg using the same #SQL and the same table names with writable cold tier.

    On August 19, 8 AM PST, catch @vyruss on the engineering + TLA+-verified distributed writes. Paul Rothrock will give a live demo.

    Q&A at the end.

    📅 us02web.zoom.us/webinar/regist

    🔗 github.com/pgEdge/ColdFront

    #DuckDB #DataEngineering #OpenSource

  7. Snowflake's pg_lake is now GA, and its data mirroring preview runs logical decoding inside Postgres via a snowflake_cdc extension rather than streaming to a consumer that knows nothing about the server.

    Better engineering. Also a tighter coupling.

    snowflake.com/en/blog/engineer

    pgexperts.com

  8. The underrated part of an AI data pipeline is the trash filter. Classification catches CTAs and promo junk before anything hits the database, and every human correction feeds back into the extraction configs. The pipeline gets better with use instead of rotting. go.upgradejs.com/dxi #AI #DataPipelines #DataEngineering

  9. Spec-Driven Data Engineering turns business rules, schemas, validation, and orchestration into versioned contracts that guide AI coding agents. hackernoon.com/why-ai-assisted #dataengineering

  10. Every CDC pipeline you run pulls a logical decoding stream from outside the server. Snowflake's new mirroring inverts it: an in-server extension (snowflake_cdc) pushes batches into Iceberg change logs plus a meta log, and the target replays that log as a state machine.

    The interesting claim is about DDL sequencing, not throughput.

    snowflake.com/en/blog/engineer

    pgexperts.com

  11. Tired of Big Tech's unpredictable cloud egress fees? 💸☁️

    While storing data in the public cloud is cheap, moving your AI datasets or Kubernetes workloads costs a fortune. The smartest move in 2026? Self-hosting your own S3-compatible Object Storage!

    We just published a step-by-step guide on deploying a secure, high-performance #MinIO server on #Ubuntu 24.04.
    servers99.com/tutorials/howto/

    #SelfHosting #Linux #DevOps #CloudComputing #SysAdmin #OpenSource #DataEngineering

  12. SapixDB calls itself the world's first agent-native database — where every data domain is owned by an AI agent, every record is cryptographically signed and hash-linked, and schemas evolve under human governance. No migrations. No DBA gates. Built from the ground up for AI agents. #Database #AI #DataEngineering

    #Database #AI #DataEngineering

  13. Whether you’re migrating to Airflow 3, scaling data pipelines, bringing AI and ML workloads into production, or defining your organization’s data strategy, you’ll leave with ideas you can put into practice.

    📅 August 31–September 2, 2026
    📍 Hyatt Regency Austin

    Don’t just watch the AI revolution unfold. Learn how to orchestrate it.

    Secure your spot: airflowsummit.org/tickets/