home.social

#apachespark — Public Fediverse posts

Live and recent posts from across the Fediverse tagged #apachespark, aggregated by home.social.

fetched live
  1. Treating SparkContext as a control tower shifts how you think about Spark: not just as an API, but as the coordinator for your entire distributed engine.

    Read More: zalt.me/blog/2026/05/sparkcont

    #ApacheSpark #SparkContext #distributed #systems

  2. AQE (Adaptive Query Execution) : adapte le plan d'exécution en temps réel

    DPP (Dynamic Partition Pruning) : ne lit que les partitions utiles pendant une jointure

    SPJ (Storage Partition Join) : évite le shuffle en utilisant le partitionnement existant

    #dataengineering #apachespark

  3. AQE (Adaptive Query Execution) : adapte le plan d'exécution en temps réel

    DPP (Dynamic Partition Pruning) : ne lit que les partitions utiles pendant une jointure

    SPJ (Storage Partition Join) : évite le shuffle en utilisant le partitionnement existant

    #dataengineering #apachespark

  4. AQE (Adaptive Query Execution) : adapte le plan d'exécution en temps réel

    DPP (Dynamic Partition Pruning) : ne lit que les partitions utiles pendant une jointure

    SPJ (Storage Partition Join) : évite le shuffle en utilisant le partitionnement existant

    #dataengineering #apachespark

  5. luminousmen.com/post/the-apach (en)

    Comment optimiser Apache Spark ?
    1. Utiliser les API DataFrame / Dataset, pas RDD.
    2. Filtrer tôt, filtrer fort.
    3. Trouver le data skew.
    4. Connaitre AQE, DPP, SPJ.
    5. Regarder l'UI.

    #dataengineering #apachespark

  6. luminousmen.com/post/the-apach (en)

    Comment optimiser Apache Spark ?
    1. Utiliser les API DataFrame / Dataset, pas RDD.
    2. Filtrer tôt, filtrer fort.
    3. Trouver le data skew.
    4. Connaitre AQE, DPP, SPJ.
    5. Regarder l'UI.

    #dataengineering #apachespark

  7. luminousmen.com/post/the-apach (en)

    Comment optimiser Apache Spark ?
    1. Utiliser les API DataFrame / Dataset, pas RDD.
    2. Filtrer tôt, filtrer fort.
    3. Trouver le data skew.
    4. Connaitre AQE, DPP, SPJ.
    5. Regarder l'UI.

    #dataengineering #apachespark

  8. 96% fewer out-of-memory (OOM) failures!

    #Pinterest shared how it improved the reliability of its #ApacheSpark workloads.

    By focusing on:
    ✅ Enhanced observability
    ✅ Configuration tuning
    ✅ Automatic memory retries

    The changes addressed persistent job failures affecting recommendation systems and large-scale data processing.

    Details here ⇨ bit.ly/4smqrQD

    #SoftwareArchitecture #BigData #CostOptimization #Memory #DistributedSystems #Observability #InfoQ

  9. 96% fewer out-of-memory (OOM) failures!

    #Pinterest shared how it improved the reliability of its #ApacheSpark workloads.

    By focusing on:
    ✅ Enhanced observability
    ✅ Configuration tuning
    ✅ Automatic memory retries

    The changes addressed persistent job failures affecting recommendation systems and large-scale data processing.

    Details here ⇨ bit.ly/4smqrQD

    #SoftwareArchitecture #BigData #CostOptimization #Memory #DistributedSystems #Observability #InfoQ

  10. 96% fewer out-of-memory (OOM) failures!

    #Pinterest shared how it improved the reliability of its #ApacheSpark workloads.

    By focusing on:
    ✅ Enhanced observability
    ✅ Configuration tuning
    ✅ Automatic memory retries

    The changes addressed persistent job failures affecting recommendation systems and large-scale data processing.

    Details here ⇨ bit.ly/4smqrQD

    #SoftwareArchitecture #BigData #CostOptimization #Memory #DistributedSystems #Observability #InfoQ

  11. 96% fewer out-of-memory (OOM) failures!

    #Pinterest shared how it improved the reliability of its #ApacheSpark workloads.

    By focusing on:
    ✅ Enhanced observability
    ✅ Configuration tuning
    ✅ Automatic memory retries

    The changes addressed persistent job failures affecting recommendation systems and large-scale data processing.

    Details here ⇨ bit.ly/4smqrQD

    #SoftwareArchitecture #BigData #CostOptimization #Memory #DistributedSystems #Observability #InfoQ

  12. 96% fewer out-of-memory (OOM) failures!

    shared how it improved the reliability of its workloads.

    By focusing on:
    ✅ Enhanced observability
    ✅ Configuration tuning
    ✅ Automatic memory retries

    The changes addressed persistent job failures affecting recommendation systems and large-scale data processing.

    Details here ⇨ bit.ly/4smqrQD

  13. Bellevue / Seattle area friends: I’m super stoked for next week’s Spark Community Spring (Friday Mar 13th: spooky 👻).

    If you’ve ever wanted to contribute to Apache Spark, come hang out and get your first Spark PR started with Felix Cheung, Huaxin Gao, Devin Petersohn, and myself :)

    We’ll help folks find starter issues, get their dev environments set up, and walk through the contribution process.

    There will be free lunch, and if enough people show up… maybe even Taco Bell for an afternoon snack*.

    #ApacheSpark #OSS #hackathon #freelunch #tacofridaymaaaaybe

    luma.com/rrfvx0ey

    (* Depends on attendance)

  14. Bellevue / Seattle area friends: I’m super stoked for next week’s Spark Community Spring (Friday Mar 13th: spooky 👻).

    If you’ve ever wanted to contribute to Apache Spark, come hang out and get your first Spark PR started with Felix Cheung, Huaxin Gao, Devin Petersohn, and myself :)

    We’ll help folks find starter issues, get their dev environments set up, and walk through the contribution process.

    There will be free lunch, and if enough people show up… maybe even Taco Bell for an afternoon snack*.

    #ApacheSpark #OSS #hackathon #freelunch #tacofridaymaaaaybe

    luma.com/rrfvx0ey

    (* Depends on attendance)

  15. Bellevue / Seattle area friends: I’m super stoked for next week’s Spark Community Spring (Friday Mar 13th: spooky 👻).

    If you’ve ever wanted to contribute to Apache Spark, come hang out and get your first Spark PR started with Felix Cheung, Huaxin Gao, Devin Petersohn, and myself :)

    We’ll help folks find starter issues, get their dev environments set up, and walk through the contribution process.

    There will be free lunch, and if enough people show up… maybe even Taco Bell for an afternoon snack*.

    #ApacheSpark #OSS #hackathon #freelunch #tacofridaymaaaaybe

    luma.com/rrfvx0ey

    (* Depends on attendance)

  16. Bellevue / Seattle area friends: I’m super stoked for next week’s Spark Community Spring (Friday Mar 13th: spooky 👻).

    If you’ve ever wanted to contribute to Apache Spark, come hang out and get your first Spark PR started with Felix Cheung, Huaxin Gao, Devin Petersohn, and myself :)

    We’ll help folks find starter issues, get their dev environments set up, and walk through the contribution process.

    There will be free lunch, and if enough people show up… maybe even Taco Bell for an afternoon snack*.

    #ApacheSpark #OSS #hackathon #freelunch #tacofridaymaaaaybe

    luma.com/rrfvx0ey

    (* Depends on attendance)

  17. Bellevue / Seattle area friends: I’m super stoked for next week’s Spark Community Spring (Friday Mar 13th: spooky 👻).

    If you’ve ever wanted to contribute to Apache Spark, come hang out and get your first Spark PR started with Felix Cheung, Huaxin Gao, Devin Petersohn, and myself :)

    We’ll help folks find starter issues, get their dev environments set up, and walk through the contribution process.

    There will be free lunch, and if enough people show up… maybe even Taco Bell for an afternoon snack*.

    #ApacheSpark #OSS #hackathon #freelunch #tacofridaymaaaaybe

    luma.com/rrfvx0ey

    (* Depends on attendance)

  18. #Pinterest launched a next-gen CDC-based ingestion framework.

    Using #ApacheKafka, #ApacheFlink, #ApacheSpark & #ApacheIceberg, they achieved:
    • Latency cut from 24+ hours to 15 minutes
    • Processing of only changed records
    • Support for incremental updates & deletions
    • Petabyte-scale data across 1,000+ pipelines

    Win: optimized cost & efficiency!

    Read the architectural deep dive on InfoQ 👉 bit.ly/4rMJB2H

    #SoftwareArchitecture #ChangeDataCapture

  19. #Pinterest launched a next-gen CDC-based ingestion framework.

    Using #ApacheKafka, #ApacheFlink, #ApacheSpark & #ApacheIceberg, they achieved:
    • Latency cut from 24+ hours to 15 minutes
    • Processing of only changed records
    • Support for incremental updates & deletions
    • Petabyte-scale data across 1,000+ pipelines

    Win: optimized cost & efficiency!

    Read the architectural deep dive on InfoQ 👉 bit.ly/4rMJB2H

    #SoftwareArchitecture #ChangeDataCapture

  20. #Pinterest launched a next-gen CDC-based ingestion framework.

    Using #ApacheKafka, #ApacheFlink, #ApacheSpark & #ApacheIceberg, they achieved:
    • Latency cut from 24+ hours to 15 minutes
    • Processing of only changed records
    • Support for incremental updates & deletions
    • Petabyte-scale data across 1,000+ pipelines

    Win: optimized cost & efficiency!

    Read the architectural deep dive on InfoQ 👉 bit.ly/4rMJB2H

    #SoftwareArchitecture #ChangeDataCapture

  21. #Pinterest launched a next-gen CDC-based ingestion framework.

    Using #ApacheKafka, #ApacheFlink, #ApacheSpark & #ApacheIceberg, they achieved:
    • Latency cut from 24+ hours to 15 minutes
    • Processing of only changed records
    • Support for incremental updates & deletions
    • Petabyte-scale data across 1,000+ pipelines

    Win: optimized cost & efficiency!

    Read the architectural deep dive on InfoQ 👉 bit.ly/4rMJB2H

    #SoftwareArchitecture #ChangeDataCapture

  22. launched a next-gen CDC-based ingestion framework.

    Using , , & , they achieved:
    • Latency cut from 24+ hours to 15 minutes
    • Processing of only changed records
    • Support for incremental updates & deletions
    • Petabyte-scale data across 1,000+ pipelines

    Win: optimized cost & efficiency!

    Read the architectural deep dive on InfoQ 👉 bit.ly/4rMJB2H

  23. In this #InfoQ article, Hina Gandhi explores a #ReinforcementLearning (RL) approach built on #ApacheSpark, enabling distributed computing systems to autonomously learn optimal configurations.

    📰 Read now: bit.ly/4r0VdyP

    #AI #bigdata #database #AIagents #InfoQ

  24. In this #InfoQ article, Hina Gandhi explores a #ReinforcementLearning (RL) approach built on #ApacheSpark, enabling distributed computing systems to autonomously learn optimal configurations.

    📰 Read now: bit.ly/4r0VdyP

    #AI #bigdata #database #AIagents #InfoQ

  25. In this #InfoQ article, Hina Gandhi explores a #ReinforcementLearning (RL) approach built on #ApacheSpark, enabling distributed computing systems to autonomously learn optimal configurations.

    📰 Read now: bit.ly/4r0VdyP

    #AI #bigdata #database #AIagents #InfoQ

  26. In this #InfoQ article, Hina Gandhi explores a #ReinforcementLearning (RL) approach built on #ApacheSpark, enabling distributed computing systems to autonomously learn optimal configurations.

    📰 Read now: bit.ly/4r0VdyP

    #AI #bigdata #database #AIagents #InfoQ

  27. In this article, Hina Gandhi explores a (RL) approach built on , enabling distributed computing systems to autonomously learn optimal configurations.

    📰 Read now: bit.ly/4r0VdyP

  28. Pinterest just shared a deep dive into Moka - its new blueprint for the future of large-scale data processing.

    The company is migrating core workloads from ageing Hadoop infrastructure to a Kubernetes-based platform on Amazon EKS, with Apache Spark as the primary engine - and support for additional frameworks coming soon.

    Curious to learn more? Read on #InfoQ 👉 bit.ly/4qCs4JP

    #DevOps #Kubernetes #AI #BigData #ApacheSpark

  29. Pinterest just shared a deep dive into Moka - its new blueprint for the future of large-scale data processing.

    The company is migrating core workloads from ageing Hadoop infrastructure to a Kubernetes-based platform on Amazon EKS, with Apache Spark as the primary engine - and support for additional frameworks coming soon.

    Curious to learn more? Read on #InfoQ 👉 bit.ly/4qCs4JP

    #DevOps #Kubernetes #AI #BigData #ApacheSpark

  30. Pinterest just shared a deep dive into Moka - its new blueprint for the future of large-scale data processing.

    The company is migrating core workloads from ageing Hadoop infrastructure to a Kubernetes-based platform on Amazon EKS, with Apache Spark as the primary engine - and support for additional frameworks coming soon.

    Curious to learn more? Read on #InfoQ 👉 bit.ly/4qCs4JP

    #DevOps #Kubernetes #AI #BigData #ApacheSpark

  31. Pinterest just shared a deep dive into Moka - its new blueprint for the future of large-scale data processing.

    The company is migrating core workloads from ageing Hadoop infrastructure to a Kubernetes-based platform on Amazon EKS, with Apache Spark as the primary engine - and support for additional frameworks coming soon.

    Curious to learn more? Read on #InfoQ 👉 bit.ly/4qCs4JP

    #DevOps #Kubernetes #AI #BigData #ApacheSpark

  32. Pinterest just shared a deep dive into Moka - its new blueprint for the future of large-scale data processing.

    The company is migrating core workloads from ageing Hadoop infrastructure to a Kubernetes-based platform on Amazon EKS, with Apache Spark as the primary engine - and support for additional frameworks coming soon.

    Curious to learn more? Read on 👉 bit.ly/4qCs4JP

  33. #CaseStudy - Agoda consolidated multiple independent data pipelines into a central #ApacheSpark platform, eliminating financial data inconsistencies.

    A multi-layered quality framework - with automated checks, ML anomaly detection, and data contracts - ensures accurate financial metrics while handling millions of daily bookings.

    Deep dive into the architecture here ⇨ bit.ly/4a109NP

    #InfoQ #SoftwareArchitecture #AI #DataPipelines

  34. #CaseStudy - Agoda consolidated multiple independent data pipelines into a central #ApacheSpark platform, eliminating financial data inconsistencies.

    A multi-layered quality framework - with automated checks, ML anomaly detection, and data contracts - ensures accurate financial metrics while handling millions of daily bookings.

    Deep dive into the architecture here ⇨ bit.ly/4a109NP

    #InfoQ #SoftwareArchitecture #AI #DataPipelines

  35. #CaseStudy - Agoda consolidated multiple independent data pipelines into a central #ApacheSpark platform, eliminating financial data inconsistencies.

    A multi-layered quality framework - with automated checks, ML anomaly detection, and data contracts - ensures accurate financial metrics while handling millions of daily bookings.

    Deep dive into the architecture here ⇨ bit.ly/4a109NP

    #InfoQ #SoftwareArchitecture #AI #DataPipelines

  36. #CaseStudy - Agoda consolidated multiple independent data pipelines into a central #ApacheSpark platform, eliminating financial data inconsistencies.

    A multi-layered quality framework - with automated checks, ML anomaly detection, and data contracts - ensures accurate financial metrics while handling millions of daily bookings.

    Deep dive into the architecture here ⇨ bit.ly/4a109NP

    #InfoQ #SoftwareArchitecture #AI #DataPipelines

  37. - Agoda consolidated multiple independent data pipelines into a central platform, eliminating financial data inconsistencies.

    A multi-layered quality framework - with automated checks, ML anomaly detection, and data contracts - ensures accurate financial metrics while handling millions of daily bookings.

    Deep dive into the architecture here ⇨ bit.ly/4a109NP

  38. Apache Spark không tự động nhanh. Tốc độ của nó phụ thuộc vào cách dùng: tránh phân vùng sai, shuffle không cần thiết và lạm dụng cache. Hiểu rõ mô hình thực thi của Spark là chìa khóa để tối ưu hiệu suất.

    #ApacheSpark #BigData #DataEngineering #Performance #Optimization #DuLieuLon #CongNgheDuLieu #HieuSuat #ToiUu

    reddit.com/r/programming/comme

  39. Apache Spark's new Declarative Pipelines framework simplifies ETL development by letting engineers focus on defining transformations while automating orchestration & error handling. This open-source solution handles batch/streaming workloads via Python/SQL interfaces, significantly reducing boilerplate code. Promising productivity gains for teams managing complex Spark pipelines. What impact might declarative approaches have on your workflow? #ApacheSpark #ETL #OpenSource

  40. Discover how Decathlon, one of the world’s leading sports retailers, adopted the #opensource library #Polars to optimize its data workflows.

    By migrating from Apache Spark to Polars for small input datasets, Decathlon achieved:
    • Significant speed
    • Meaningful cost savings

    👉 Learn more: bit.ly/4qmb2zc

    #InfoQ #AI #ApacheSpark