home.social

#dataengineering — Public Fediverse posts

Live and recent posts from across the Fediverse tagged #dataengineering, aggregated by home.social.

fetched live
  1. Every transformation tested, every model documented, lineage you can trace end to end. My dbt project treats analytics like software.
    #dbt #DataEngineering #Analytics #SQL

  2. Every transformation tested, every model documented, lineage you can trace end to end. My dbt project treats analytics like software.
    #dbt #DataEngineering #Analytics #SQL

  3. Every transformation tested, every model documented, lineage you can trace end to end. My dbt project treats analytics like software.
    #dbt #DataEngineering #Analytics #SQL

  4. Every transformation tested, every model documented, lineage you can trace end to end. My dbt project treats analytics like software.
    #dbt #DataEngineering #Analytics #SQL

  5. Zero data loss in migrations isn’t a promise—it’s an engineering outcome. Learn how mapping, validation, exception handling, and auditability make it provable. hackernoon.com/zero-data-loss- #dataengineering

  6. Zero data loss in migrations isn’t a promise—it’s an engineering outcome. Learn how mapping, validation, exception handling, and auditability make it provable. hackernoon.com/zero-data-loss- #dataengineering

  7. Zero data loss in migrations isn’t a promise—it’s an engineering outcome. Learn how mapping, validation, exception handling, and auditability make it provable. hackernoon.com/zero-data-loss- #dataengineering

  8. Zero data loss in migrations isn’t a promise—it’s an engineering outcome. Learn how mapping, validation, exception handling, and auditability make it provable. hackernoon.com/zero-data-loss-

  9. Zero data loss in migrations isn’t a promise—it’s an engineering outcome. Learn how mapping, validation, exception handling, and auditability make it provable. hackernoon.com/zero-data-loss- #dataengineering

  10. 🐍PyBay 2026 Speaker Highlight🎤 Rodrigo Silva Ferreira

    Building scientific observatories with Python.

    Rodrigo Silva Ferreira is a developer tooling and AI practitioner with a background in scientific computing and software quality.

    📍 Oct. 03, 2026, San Francisco, CA: pybay.org/
    🎤 More talks: pybay.org/speaking/speaker-pro
    🎟️ Get your tickets at: pybay.org/get-tickets/

    #PyBay #Python #PyBay2026 #Data #DataEngineering #DataVisualization #PythonForScience

  11. 🐍PyBay 2026 Speaker Highlight🎤 Rodrigo Silva Ferreira

    Building scientific observatories with Python.

    Rodrigo Silva Ferreira is a developer tooling and AI practitioner with a background in scientific computing and software quality.

    📍 Oct. 03, 2026, San Francisco, CA: pybay.org/
    🎤 More talks: pybay.org/speaking/speaker-pro
    🎟️ Get your tickets at: pybay.org/get-tickets/

  12. 🐍PyBay 2026 Speaker Highlight🎤 Rodrigo Silva Ferreira

    Building scientific observatories with Python.

    Rodrigo Silva Ferreira is a developer tooling and AI practitioner with a background in scientific computing and software quality.

    📍 Oct. 03, 2026, San Francisco, CA: pybay.org/
    🎤 More talks: pybay.org/speaking/speaker-pro
    🎟️ Get your tickets at: pybay.org/get-tickets/

    #PyBay #Python #PyBay2026 #Data #DataEngineering #DataVisualization #PythonForScience

  13. 🐍PyBay 2026 Speaker Highlight🎤 Rodrigo Silva Ferreira

    Building scientific observatories with Python.

    Rodrigo Silva Ferreira is a developer tooling and AI practitioner with a background in scientific computing and software quality.

    📍 Oct. 03, 2026, San Francisco, CA: pybay.org/
    🎤 More talks: pybay.org/speaking/speaker-pro
    🎟️ Get your tickets at: pybay.org/get-tickets/

    #PyBay #Python #PyBay2026 #Data #DataEngineering #DataVisualization #PythonForScience

  14. 🐍PyBay 2026 Speaker Highlight🎤 Rodrigo Silva Ferreira

    Building scientific observatories with Python.

    Rodrigo Silva Ferreira is a developer tooling and AI practitioner with a background in scientific computing and software quality.

    📍 Oct. 03, 2026, San Francisco, CA: pybay.org/
    🎤 More talks: pybay.org/speaking/speaker-pro
    🎟️ Get your tickets at: pybay.org/get-tickets/

    #PyBay #Python #PyBay2026 #Data #DataEngineering #DataVisualization #PythonForScience

  15. AI staff augmentation in Miami gives tech leaders a direct route to senior ML and data engineering talent without the multi-month recruiting timeline. Plug verified specialists into your existing team and infrastructure. #AI #Miami #StaffAugmentation #MachineLearning #DataEngineering

    Explore how we can help at codeponents.com/ai-staff-augme

  16. AI staff augmentation in Miami gives tech leaders a direct route to senior ML and data engineering talent without the multi-month recruiting timeline. Plug verified specialists into your existing team and infrastructure. #AI #Miami #StaffAugmentation #MachineLearning #DataEngineering

    Explore how we can help at codeponents.com/ai-staff-augme

  17. AI staff augmentation in Miami gives tech leaders a direct route to senior ML and data engineering talent without the multi-month recruiting timeline. Plug verified specialists into your existing team and infrastructure. #AI #Miami #StaffAugmentation #MachineLearning #DataEngineering

    Explore how we can help at codeponents.com/ai-staff-augme

  18. AI staff augmentation in Miami gives tech leaders a direct route to senior ML and data engineering talent without the multi-month recruiting timeline. Plug verified specialists into your existing team and infrastructure. #AI #Miami #StaffAugmentation #MachineLearning #DataEngineering

    Explore how we can help at codeponents.com/ai-staff-augme

  19. AI staff augmentation in Miami gives tech leaders a direct route to senior ML and data engineering talent without the multi-month recruiting timeline. Plug verified specialists into your existing team and infrastructure. #AI #Miami #StaffAugmentation #MachineLearning #DataEngineering

    Explore how we can help at codeponents.com/ai-staff-augme

  20. We have big news to share! 🎉

    AWS has become a Platinum Sponsor of Psycopg!

    For two decades, Psycopg has been the standard way Python talks to PostgreSQL. This sponsorship reflects the role Psycopg plays across the Python and PostgreSQL ecosystem, and will fund the ongoing development of Psycopg 3 and 2.

    If your organisation depends on Psycopg, consider supporting us. Links in replies 👇

    #PostgreSQL #Python #Psycopg #AWS #OpenSource #FreeSoftware #FOSS #DataEngineering #SponsorOpenSource

  21. We have big news to share! 🎉

    AWS has become a Platinum Sponsor of Psycopg!

    For two decades, Psycopg has been the standard way Python talks to PostgreSQL. This sponsorship reflects the role Psycopg plays across the Python and PostgreSQL ecosystem, and will fund the ongoing development of Psycopg 3 and 2.

    If your organisation depends on Psycopg, consider supporting us. Links in replies 👇

  22. We have big news to share! 🎉

    AWS has become a Platinum Sponsor of Psycopg!

    For two decades, Psycopg has been the standard way Python talks to PostgreSQL. This sponsorship reflects the role Psycopg plays across the Python and PostgreSQL ecosystem, and will fund the ongoing development of Psycopg 3 and 2.

    If your organisation depends on Psycopg, consider supporting us. Links in replies 👇

    #PostgreSQL #Python #Psycopg #AWS #OpenSource #FreeSoftware #FOSS #DataEngineering #SponsorOpenSource

  23. We have big news to share! 🎉

    AWS has become a Platinum Sponsor of Psycopg!

    For two decades, Psycopg has been the standard way Python talks to PostgreSQL. This sponsorship reflects the role Psycopg plays across the Python and PostgreSQL ecosystem, and will fund the ongoing development of Psycopg 3 and 2.

    If your organisation depends on Psycopg, consider supporting us. Links in replies 👇

    #PostgreSQL #Python #Psycopg #AWS #OpenSource #FreeSoftware #FOSS #DataEngineering #SponsorOpenSource

  24. Billions of events, one streaming pipeline: Kafka into Spark, aggregations out to a live dashboard. One of eight big-data projects scoped to genuinely large public data.
    #DataEngineering #Kafka #Spark #BigData

  25. Billions of events, one streaming pipeline: Kafka into Spark, aggregations out to a live dashboard. One of eight big-data projects scoped to genuinely large public data.
    #DataEngineering #Kafka #Spark #BigData

  26. Billions of events, one streaming pipeline: Kafka into Spark, aggregations out to a live dashboard. One of eight big-data projects scoped to genuinely large public data.
    #DataEngineering #Kafka #Spark #BigData

  27. Агент пишет код. Что тогда остаётся инженеру данных?

    В работе инженера данных хватает повторяющихся задач: забрать данные из API, проверить структуру, преобразовать поля, записать результат в хранилище. Вокруг этого — настройки, тесты, сообщения об ошибках и документация. Такую работу удобно поручать Агенту. Опять агенты

    habr.com/ru/articles/1084096/

    #сезон_код_будущего #ai #aiагенты #agents #dataengineering #dataengineer #data #python #ии #ииассистент

  28. Агент пишет код. Что тогда остаётся инженеру данных?

    В работе инженера данных хватает повторяющихся задач: забрать данные из API, проверить структуру, преобразовать поля, записать результат в хранилище. Вокруг этого — настройки, тесты, сообщения об ошибках и документация. Такую работу удобно поручать Агенту. Опять агенты

    habr.com/ru/articles/1084096/

    #сезон_код_будущего #ai #aiагенты #agents #dataengineering #dataengineer #data #python #ии #ииассистент

  29. Managing and visualizing thousands of spatial data points per second requires specific, highly scalable infrastructure. Here is a technical breakdown of an open-source GIS architecture designed specifically to ingest, process, and map real-time IoT sensor data and telemetry with minimal latency.

    Click here: rsandgis.me/articles/open-sour

    #IoT #GIS #Architecture #DataEngineering #SpatialData #RealTime

  30. Managing and visualizing thousands of spatial data points per second requires specific, highly scalable infrastructure. Here is a technical breakdown of an open-source GIS architecture designed specifically to ingest, process, and map real-time IoT sensor data and telemetry with minimal latency.

    Click here: rsandgis.me/articles/open-sour

    #IoT #GIS #Architecture #DataEngineering #SpatialData #RealTime

  31. Managing and visualizing thousands of spatial data points per second requires specific, highly scalable infrastructure. Here is a technical breakdown of an open-source GIS architecture designed specifically to ingest, process, and map real-time IoT sensor data and telemetry with minimal latency.

    Click here: rsandgis.me/articles/open-sour

    #IoT #GIS #Architecture #DataEngineering #SpatialData #RealTime

  32. Managing and visualizing thousands of spatial data points per second requires specific, highly scalable infrastructure. Here is a technical breakdown of an open-source GIS architecture designed specifically to ingest, process, and map real-time IoT sensor data and telemetry with minimal latency.

    Click here: rsandgis.me/articles/open-sour

    #IoT #GIS #Architecture #DataEngineering #SpatialData #RealTime

  33. Struggling with painfully slow spatial queries on large vector layers? Learn advanced indexing strategies, spatial clustering techniques, and critical PostgreSQL configuration tweaks to maximize the read and write performance of your open-source spatial database architectures.

    Click here: rsandgis.me/articles/open-sour

    #PostGIS #Database #Geospatial #DataEngineering #GIS #PostgreSQL

  34. Struggling with painfully slow spatial queries on large vector layers? Learn advanced indexing strategies, spatial clustering techniques, and critical PostgreSQL configuration tweaks to maximize the read and write performance of your open-source spatial database architectures.

    Click here: rsandgis.me/articles/open-sour

    #PostGIS #Database #Geospatial #DataEngineering #GIS #PostgreSQL

  35. Struggling with painfully slow spatial queries on large vector layers? Learn advanced indexing strategies, spatial clustering techniques, and critical PostgreSQL configuration tweaks to maximize the read and write performance of your open-source spatial database architectures.

    Click here: rsandgis.me/articles/open-sour

    #PostGIS #Database #Geospatial #DataEngineering #GIS #PostgreSQL

  36. 🚀✨ Congrats, techies! You've finally found a way to make #ClickHouse and #Postgres talk faster than a caffeinated squirrel 🐿️. Next up: convincing your boss why they should care about this monumental achievement in data yak-shaving. 🙄🔧
    clickhouse.com/blog/introducin #DataEngineering #TechAchievement #DataYakShaving #HackerNews #ngated

  37. 🚀✨ Congrats, techies! You've finally found a way to make #ClickHouse and #Postgres talk faster than a caffeinated squirrel 🐿️. Next up: convincing your boss why they should care about this monumental achievement in data yak-shaving. 🙄🔧
    clickhouse.com/blog/introducin #DataEngineering #TechAchievement #DataYakShaving #HackerNews #ngated

  38. 🚀✨ Congrats, techies! You've finally found a way to make #ClickHouse and #Postgres talk faster than a caffeinated squirrel 🐿️. Next up: convincing your boss why they should care about this monumental achievement in data yak-shaving. 🙄🔧
    clickhouse.com/blog/introducin #DataEngineering #TechAchievement #DataYakShaving #HackerNews #ngated

  39. 🚀✨ Congrats, techies! You've finally found a way to make #ClickHouse and #Postgres talk faster than a caffeinated squirrel 🐿️. Next up: convincing your boss why they should care about this monumental achievement in data yak-shaving. 🙄🔧
    clickhouse.com/blog/introducin #DataEngineering #TechAchievement #DataYakShaving #HackerNews #ngated

  40. Zax looks like a very cool new tool. Bringing a SQL interface to high-dimensional array data in the cloud. Designed to bridge the gap between monster scientific datasets and the often tabular slices and aggregations that get used downstream.

    earthmover.io/blog/compute-roa

    #SQL #DataEngineering #Zarr #Xarray

  41. Zax looks like a very cool new tool. Bringing a SQL interface to high-dimensional array data in the cloud. Designed to bridge the gap between monster scientific datasets and the often tabular slices and aggregations that get used downstream.

    earthmover.io/blog/compute-roa

    #SQL #DataEngineering #Zarr #Xarray

  42. Zax looks like a very cool new tool. Bringing a SQL interface to high-dimensional array data in the cloud. Designed to bridge the gap between monster scientific datasets and the often tabular slices and aggregations that get used downstream.

    earthmover.io/blog/compute-roa

    #SQL #DataEngineering #Zarr #Xarray

  43. Zax looks like a very cool new tool. Bringing a SQL interface to high-dimensional array data in the cloud. Designed to bridge the gap between monster scientific datasets and the often tabular slices and aggregations that get used downstream.

    earthmover.io/blog/compute-roa

    #SQL #DataEngineering #Zarr #Xarray

  44. 180 lines of Python using great-expectations to validate datasets against YAML rule sets and output reports. For data engineers who want checks in CI without building a valtersit.com/python/data-qual #python #dataengineering #greatexpectations

  45. 180 lines of Python using great-expectations to validate datasets against YAML rule sets and output reports. For data engineers who want checks in CI without building a valtersit.com/python/data-qual #python #dataengineering #greatexpectations

  46. 180 lines of Python using great-expectations to validate datasets against YAML rule sets and output reports. For data engineers who want checks in CI without building a valtersit.com/python/data-qual #python #dataengineering #greatexpectations

  47. The real problem with enterprise data? Not multiple definitions - pretending there’s only one.

    If Marketing and Finance can’t agree on what an “active customer” is, how can an AI agent?

    Enterprise AI agents don’t fail because models are dumb. They fail because enterprise semantics are messy!

    🔗 Watch Fabiane Nardon’s full #QConAI Boston talk: infoq.com/presentations/enterp

    #AI #DataEngineering #EnterpriseAI #AIAgents #InfoQ

  48. The real problem with enterprise data? Not multiple definitions - pretending there’s only one.

    If Marketing and Finance can’t agree on what an “active customer” is, how can an AI agent?

    Enterprise AI agents don’t fail because models are dumb. They fail because enterprise semantics are messy!

    🔗 Watch Fabiane Nardon’s full #QConAI Boston talk: infoq.com/presentations/enterp

    #AI #DataEngineering #EnterpriseAI #AIAgents #InfoQ

  49. The real problem with enterprise data? Not multiple definitions - pretending there’s only one.

    If Marketing and Finance can’t agree on what an “active customer” is, how can an AI agent?

    Enterprise AI agents don’t fail because models are dumb. They fail because enterprise semantics are messy!

    🔗 Watch Fabiane Nardon’s full #QConAI Boston talk: infoq.com/presentations/enterp

    #AI #DataEngineering #EnterpriseAI #AIAgents #InfoQ

  50. The real problem with enterprise data? Not multiple definitions - pretending there’s only one.

    If Marketing and Finance can’t agree on what an “active customer” is, how can an AI agent?

    Enterprise AI agents don’t fail because models are dumb. They fail because enterprise semantics are messy!

    🔗 Watch Fabiane Nardon’s full Boston talk: infoq.com/presentations/enterp

  51. Stop wrestling with CSV bloat. New script converts CSVs to compressed Parquet with PyArrow, squeezing storage and speeding up your pipelines. Intermediate level, 180 lines. Try it: valtersit.com/python/parquet-f #Python #DataEngineering

  52. Stop wrestling with CSV bloat. New script converts CSVs to compressed Parquet with PyArrow, squeezing storage and speeding up your pipelines. Intermediate level, 180 lines. Try it: valtersit.com/python/parquet-f #Python #DataEngineering

  53. Stop wrestling with CSV bloat. New script converts CSVs to compressed Parquet with PyArrow, squeezing storage and speeding up your pipelines. Intermediate level, 180 lines. Try it: valtersit.com/python/parquet-f #Python #DataEngineering