#apachespark — Public Fediverse posts
Live and recent posts from across the Fediverse tagged #apachespark, aggregated by home.social.
-
Stackable Data Platform 26.7: Security, SBOMs, and Flexible Registries
Stackable SDP 26.7 focuses on quality and supply chain security, bringing SLSA provenance and breaking changes to registry logic.
#ApacheSpark #Containerisierung #IT #Kubernetes #OpenSource #Updates #news
-
Stackable Data Platform 26.7: Security, SBOMs, and Flexible Registries
Stackable SDP 26.7 focuses on quality and supply chain security, bringing SLSA provenance and breaking changes to registry logic.
#ApacheSpark #Containerisierung #IT #Kubernetes #OpenSource #Updates #news
-
Stackable Data Platform 26.7: Security, SBOMs, and Flexible Registries
Stackable SDP 26.7 focuses on quality and supply chain security, bringing SLSA provenance and breaking changes to registry logic.
#ApacheSpark #Containerisierung #IT #Kubernetes #OpenSource #Updates #news
-
Stackable Data Platform 26.7: Security, SBOMs, and Flexible Registries
Stackable SDP 26.7 focuses on quality and supply chain security, bringing SLSA provenance and breaking changes to registry logic.
#ApacheSpark #Containerisierung #IT #Kubernetes #OpenSource #Updates #news
-
Stackable Data Platform 26.7: Security, SBOMs, and Flexible Registries
Stackable SDP 26.7 focuses on quality and supply chain security, bringing SLSA provenance and breaking changes to registry logic.
#ApacheSpark #Containerisierung #IT #Kubernetes #OpenSource #Updates #news
-
Stackable Data Platform 26.7: Sicherheit, SBOMs und flexible Registries
Stackable stellt die Data Platform 26.7 auf Qualität und Supply-Chain-Sicherheit um. Neben SLSA-Provenance drohen bei der Registry-Logik Breaking Changes.
#ApacheSpark #Containerisierung #IT #Kubernetes #OpenSource #Updates #news
-
Stackable Data Platform 26.7: Sicherheit, SBOMs und flexible Registries
Stackable stellt die Data Platform 26.7 auf Qualität und Supply-Chain-Sicherheit um. Neben SLSA-Provenance drohen bei der Registry-Logik Breaking Changes.
#ApacheSpark #Containerisierung #IT #Kubernetes #OpenSource #Updates #news
-
Stackable Data Platform 26.7: Sicherheit, SBOMs und flexible Registries
Stackable stellt die Data Platform 26.7 auf Qualität und Supply-Chain-Sicherheit um. Neben SLSA-Provenance drohen bei der Registry-Logik Breaking Changes.
#ApacheSpark #Containerisierung #IT #Kubernetes #OpenSource #Updates #news
-
Stackable Data Platform 26.7: Sicherheit, SBOMs und flexible Registries
Stackable stellt die Data Platform 26.7 auf Qualität und Supply-Chain-Sicherheit um. Neben SLSA-Provenance drohen bei der Registry-Logik Breaking Changes.
#ApacheSpark #Containerisierung #IT #Kubernetes #OpenSource #Updates #news
-
Stackable Data Platform 26.7: Sicherheit, SBOMs und flexible Registries
Stackable stellt die Data Platform 26.7 auf Qualität und Supply-Chain-Sicherheit um. Neben SLSA-Provenance drohen bei der Registry-Logik Breaking Changes.
#ApacheSpark #Containerisierung #IT #Kubernetes #OpenSource #Updates #news
-
Treating SparkContext as a control tower shifts how you think about Spark: not just as an API, but as the coordinator for your entire distributed engine.
Read More: https://zalt.me/blog/2026/05/sparkcontext-control-tower
-
AQE (Adaptive Query Execution) : adapte le plan d'exécution en temps réel
DPP (Dynamic Partition Pruning) : ne lit que les partitions utiles pendant une jointure
SPJ (Storage Partition Join) : évite le shuffle en utilisant le partitionnement existant
-
AQE (Adaptive Query Execution) : adapte le plan d'exécution en temps réel
DPP (Dynamic Partition Pruning) : ne lit que les partitions utiles pendant une jointure
SPJ (Storage Partition Join) : évite le shuffle en utilisant le partitionnement existant
-
AQE (Adaptive Query Execution) : adapte le plan d'exécution en temps réel
DPP (Dynamic Partition Pruning) : ne lit que les partitions utiles pendant une jointure
SPJ (Storage Partition Join) : évite le shuffle en utilisant le partitionnement existant
-
https://luminousmen.com/post/the-apache-spark-optimization-checklist/ (en)
Comment optimiser Apache Spark ?
1. Utiliser les API DataFrame / Dataset, pas RDD.
2. Filtrer tôt, filtrer fort.
3. Trouver le data skew.
4. Connaitre AQE, DPP, SPJ.
5. Regarder l'UI. -
https://luminousmen.com/post/the-apache-spark-optimization-checklist/ (en)
Comment optimiser Apache Spark ?
1. Utiliser les API DataFrame / Dataset, pas RDD.
2. Filtrer tôt, filtrer fort.
3. Trouver le data skew.
4. Connaitre AQE, DPP, SPJ.
5. Regarder l'UI. -
https://luminousmen.com/post/the-apache-spark-optimization-checklist/ (en)
Comment optimiser Apache Spark ?
1. Utiliser les API DataFrame / Dataset, pas RDD.
2. Filtrer tôt, filtrer fort.
3. Trouver le data skew.
4. Connaitre AQE, DPP, SPJ.
5. Regarder l'UI. -
https://www.europesays.com/ie/467833/ From Batch to Micro-Batch Streaming: Lessons Learned the Hard Way in a Delta Index Pipeline #AI #ApacheSpark #BigData #database #Development #Éire #IE #Infrastructure #Ireland #MicroBatchStreamingLessonsLearned #ML&DataEngineering #SparkStreaming #Streaming #Technology
-
Join me Tuesday for my next Python Data Science & AI Full Throttle! https://deitel.com/PYDSFT
O'Reilly Media Pearson #deitel #python #machinelearning #deeplearning #NLP #datamining #ApacheSpark #BigData #IoT #GenAI
-
Join me Tuesday for my next Python Data Science & AI Full Throttle! https://deitel.com/PYDSFT
O'Reilly Media Pearson #deitel #python #machinelearning #deeplearning #NLP #datamining #ApacheSpark #BigData #IoT #GenAI
-
Join me Tuesday for my next Python Data Science & AI Full Throttle! https://deitel.com/PYDSFT
O'Reilly Media Pearson #deitel #python #machinelearning #deeplearning #NLP #datamining #ApacheSpark #BigData #IoT #GenAI
-
Join me Tuesday for my next Python Data Science & AI Full Throttle! https://deitel.com/PYDSFT
O'Reilly Media Pearson #deitel #python #machinelearning #deeplearning #NLP #datamining #ApacheSpark #BigData #IoT #GenAI
-
Join me Tuesday for my next Python Data Science & AI Full Throttle! https://deitel.com/PYDSFT
O'Reilly Media Pearson #deitel #python #machinelearning #deeplearning #NLP #datamining #ApacheSpark #BigData #IoT #GenAI
-
96% fewer out-of-memory (OOM) failures!
#Pinterest shared how it improved the reliability of its #ApacheSpark workloads.
By focusing on:
✅ Enhanced observability
✅ Configuration tuning
✅ Automatic memory retriesThe changes addressed persistent job failures affecting recommendation systems and large-scale data processing.
Details here ⇨ https://bit.ly/4smqrQD
#SoftwareArchitecture #BigData #CostOptimization #Memory #DistributedSystems #Observability #InfoQ
-
96% fewer out-of-memory (OOM) failures!
#Pinterest shared how it improved the reliability of its #ApacheSpark workloads.
By focusing on:
✅ Enhanced observability
✅ Configuration tuning
✅ Automatic memory retriesThe changes addressed persistent job failures affecting recommendation systems and large-scale data processing.
Details here ⇨ https://bit.ly/4smqrQD
#SoftwareArchitecture #BigData #CostOptimization #Memory #DistributedSystems #Observability #InfoQ
-
96% fewer out-of-memory (OOM) failures!
#Pinterest shared how it improved the reliability of its #ApacheSpark workloads.
By focusing on:
✅ Enhanced observability
✅ Configuration tuning
✅ Automatic memory retriesThe changes addressed persistent job failures affecting recommendation systems and large-scale data processing.
Details here ⇨ https://bit.ly/4smqrQD
#SoftwareArchitecture #BigData #CostOptimization #Memory #DistributedSystems #Observability #InfoQ
-
96% fewer out-of-memory (OOM) failures!
#Pinterest shared how it improved the reliability of its #ApacheSpark workloads.
By focusing on:
✅ Enhanced observability
✅ Configuration tuning
✅ Automatic memory retriesThe changes addressed persistent job failures affecting recommendation systems and large-scale data processing.
Details here ⇨ https://bit.ly/4smqrQD
#SoftwareArchitecture #BigData #CostOptimization #Memory #DistributedSystems #Observability #InfoQ
-
96% fewer out-of-memory (OOM) failures!
#Pinterest shared how it improved the reliability of its #ApacheSpark workloads.
By focusing on:
✅ Enhanced observability
✅ Configuration tuning
✅ Automatic memory retriesThe changes addressed persistent job failures affecting recommendation systems and large-scale data processing.
Details here ⇨ https://bit.ly/4smqrQD
#SoftwareArchitecture #BigData #CostOptimization #Memory #DistributedSystems #Observability #InfoQ
-
The Data Lakehouse Explained: Why Apache Iceberg Is Quietly Running the Show
https://techlife.blog/posts/data-lakehouse-iceberg
#ApacheIceberg #DataLakehouse #DataWarehouse #DataLake #Snowflake #ApacheSpark #DataEngineering
-
The Data Lakehouse Explained: Why Apache Iceberg Is Quietly Running the Show
https://techlife.blog/posts/data-lakehouse-iceberg
#ApacheIceberg #DataLakehouse #DataWarehouse #DataLake #Snowflake #ApacheSpark #DataEngineering
-
Bellevue / Seattle area friends: I’m super stoked for next week’s Spark Community Spring (Friday Mar 13th: spooky 👻).
If you’ve ever wanted to contribute to Apache Spark, come hang out and get your first Spark PR started with Felix Cheung, Huaxin Gao, Devin Petersohn, and myself :)
We’ll help folks find starter issues, get their dev environments set up, and walk through the contribution process.
There will be free lunch, and if enough people show up… maybe even Taco Bell for an afternoon snack*.
#ApacheSpark #OSS #hackathon #freelunch #tacofridaymaaaaybe
(* Depends on attendance)
-
Bellevue / Seattle area friends: I’m super stoked for next week’s Spark Community Spring (Friday Mar 13th: spooky 👻).
If you’ve ever wanted to contribute to Apache Spark, come hang out and get your first Spark PR started with Felix Cheung, Huaxin Gao, Devin Petersohn, and myself :)
We’ll help folks find starter issues, get their dev environments set up, and walk through the contribution process.
There will be free lunch, and if enough people show up… maybe even Taco Bell for an afternoon snack*.
#ApacheSpark #OSS #hackathon #freelunch #tacofridaymaaaaybe
(* Depends on attendance)
-
Bellevue / Seattle area friends: I’m super stoked for next week’s Spark Community Spring (Friday Mar 13th: spooky 👻).
If you’ve ever wanted to contribute to Apache Spark, come hang out and get your first Spark PR started with Felix Cheung, Huaxin Gao, Devin Petersohn, and myself :)
We’ll help folks find starter issues, get their dev environments set up, and walk through the contribution process.
There will be free lunch, and if enough people show up… maybe even Taco Bell for an afternoon snack*.
#ApacheSpark #OSS #hackathon #freelunch #tacofridaymaaaaybe
(* Depends on attendance)
-
Bellevue / Seattle area friends: I’m super stoked for next week’s Spark Community Spring (Friday Mar 13th: spooky 👻).
If you’ve ever wanted to contribute to Apache Spark, come hang out and get your first Spark PR started with Felix Cheung, Huaxin Gao, Devin Petersohn, and myself :)
We’ll help folks find starter issues, get their dev environments set up, and walk through the contribution process.
There will be free lunch, and if enough people show up… maybe even Taco Bell for an afternoon snack*.
#ApacheSpark #OSS #hackathon #freelunch #tacofridaymaaaaybe
(* Depends on attendance)
-
Bellevue / Seattle area friends: I’m super stoked for next week’s Spark Community Spring (Friday Mar 13th: spooky 👻).
If you’ve ever wanted to contribute to Apache Spark, come hang out and get your first Spark PR started with Felix Cheung, Huaxin Gao, Devin Petersohn, and myself :)
We’ll help folks find starter issues, get their dev environments set up, and walk through the contribution process.
There will be free lunch, and if enough people show up… maybe even Taco Bell for an afternoon snack*.
#ApacheSpark #OSS #hackathon #freelunch #tacofridaymaaaaybe
(* Depends on attendance)
-
#Pinterest launched a next-gen CDC-based ingestion framework.
Using #ApacheKafka, #ApacheFlink, #ApacheSpark & #ApacheIceberg, they achieved:
• Latency cut from 24+ hours to 15 minutes
• Processing of only changed records
• Support for incremental updates & deletions
• Petabyte-scale data across 1,000+ pipelinesWin: optimized cost & efficiency!
Read the architectural deep dive on InfoQ 👉 https://bit.ly/4rMJB2H
-
#Pinterest launched a next-gen CDC-based ingestion framework.
Using #ApacheKafka, #ApacheFlink, #ApacheSpark & #ApacheIceberg, they achieved:
• Latency cut from 24+ hours to 15 minutes
• Processing of only changed records
• Support for incremental updates & deletions
• Petabyte-scale data across 1,000+ pipelinesWin: optimized cost & efficiency!
Read the architectural deep dive on InfoQ 👉 https://bit.ly/4rMJB2H
-
#Pinterest launched a next-gen CDC-based ingestion framework.
Using #ApacheKafka, #ApacheFlink, #ApacheSpark & #ApacheIceberg, they achieved:
• Latency cut from 24+ hours to 15 minutes
• Processing of only changed records
• Support for incremental updates & deletions
• Petabyte-scale data across 1,000+ pipelinesWin: optimized cost & efficiency!
Read the architectural deep dive on InfoQ 👉 https://bit.ly/4rMJB2H
-
#Pinterest launched a next-gen CDC-based ingestion framework.
Using #ApacheKafka, #ApacheFlink, #ApacheSpark & #ApacheIceberg, they achieved:
• Latency cut from 24+ hours to 15 minutes
• Processing of only changed records
• Support for incremental updates & deletions
• Petabyte-scale data across 1,000+ pipelinesWin: optimized cost & efficiency!
Read the architectural deep dive on InfoQ 👉 https://bit.ly/4rMJB2H
-
#Pinterest launched a next-gen CDC-based ingestion framework.
Using #ApacheKafka, #ApacheFlink, #ApacheSpark & #ApacheIceberg, they achieved:
• Latency cut from 24+ hours to 15 minutes
• Processing of only changed records
• Support for incremental updates & deletions
• Petabyte-scale data across 1,000+ pipelinesWin: optimized cost & efficiency!
Read the architectural deep dive on InfoQ 👉 https://bit.ly/4rMJB2H
-
In this #InfoQ article, Hina Gandhi explores a #ReinforcementLearning (RL) approach built on #ApacheSpark, enabling distributed computing systems to autonomously learn optimal configurations.
📰 Read now: https://bit.ly/4r0VdyP
-
In this #InfoQ article, Hina Gandhi explores a #ReinforcementLearning (RL) approach built on #ApacheSpark, enabling distributed computing systems to autonomously learn optimal configurations.
📰 Read now: https://bit.ly/4r0VdyP
-
In this #InfoQ article, Hina Gandhi explores a #ReinforcementLearning (RL) approach built on #ApacheSpark, enabling distributed computing systems to autonomously learn optimal configurations.
📰 Read now: https://bit.ly/4r0VdyP
-
In this #InfoQ article, Hina Gandhi explores a #ReinforcementLearning (RL) approach built on #ApacheSpark, enabling distributed computing systems to autonomously learn optimal configurations.
📰 Read now: https://bit.ly/4r0VdyP
-
In this #InfoQ article, Hina Gandhi explores a #ReinforcementLearning (RL) approach built on #ApacheSpark, enabling distributed computing systems to autonomously learn optimal configurations.
📰 Read now: https://bit.ly/4r0VdyP
-
Pinterest just shared a deep dive into Moka - its new blueprint for the future of large-scale data processing.
The company is migrating core workloads from ageing Hadoop infrastructure to a Kubernetes-based platform on Amazon EKS, with Apache Spark as the primary engine - and support for additional frameworks coming soon.
Curious to learn more? Read on #InfoQ 👉 https://bit.ly/4qCs4JP
-
Pinterest just shared a deep dive into Moka - its new blueprint for the future of large-scale data processing.
The company is migrating core workloads from ageing Hadoop infrastructure to a Kubernetes-based platform on Amazon EKS, with Apache Spark as the primary engine - and support for additional frameworks coming soon.
Curious to learn more? Read on #InfoQ 👉 https://bit.ly/4qCs4JP
-
Pinterest just shared a deep dive into Moka - its new blueprint for the future of large-scale data processing.
The company is migrating core workloads from ageing Hadoop infrastructure to a Kubernetes-based platform on Amazon EKS, with Apache Spark as the primary engine - and support for additional frameworks coming soon.
Curious to learn more? Read on #InfoQ 👉 https://bit.ly/4qCs4JP
-
Pinterest just shared a deep dive into Moka - its new blueprint for the future of large-scale data processing.
The company is migrating core workloads from ageing Hadoop infrastructure to a Kubernetes-based platform on Amazon EKS, with Apache Spark as the primary engine - and support for additional frameworks coming soon.
Curious to learn more? Read on #InfoQ 👉 https://bit.ly/4qCs4JP
-
Pinterest just shared a deep dive into Moka - its new blueprint for the future of large-scale data processing.
The company is migrating core workloads from ageing Hadoop infrastructure to a Kubernetes-based platform on Amazon EKS, with Apache Spark as the primary engine - and support for additional frameworks coming soon.
Curious to learn more? Read on #InfoQ 👉 https://bit.ly/4qCs4JP
-
#CaseStudy - Agoda consolidated multiple independent data pipelines into a central #ApacheSpark platform, eliminating financial data inconsistencies.
A multi-layered quality framework - with automated checks, ML anomaly detection, and data contracts - ensures accurate financial metrics while handling millions of daily bookings.
Deep dive into the architecture here ⇨ https://bit.ly/4a109NP
-
#CaseStudy - Agoda consolidated multiple independent data pipelines into a central #ApacheSpark platform, eliminating financial data inconsistencies.
A multi-layered quality framework - with automated checks, ML anomaly detection, and data contracts - ensures accurate financial metrics while handling millions of daily bookings.
Deep dive into the architecture here ⇨ https://bit.ly/4a109NP
-
#CaseStudy - Agoda consolidated multiple independent data pipelines into a central #ApacheSpark platform, eliminating financial data inconsistencies.
A multi-layered quality framework - with automated checks, ML anomaly detection, and data contracts - ensures accurate financial metrics while handling millions of daily bookings.
Deep dive into the architecture here ⇨ https://bit.ly/4a109NP
-
#CaseStudy - Agoda consolidated multiple independent data pipelines into a central #ApacheSpark platform, eliminating financial data inconsistencies.
A multi-layered quality framework - with automated checks, ML anomaly detection, and data contracts - ensures accurate financial metrics while handling millions of daily bookings.
Deep dive into the architecture here ⇨ https://bit.ly/4a109NP
-
#CaseStudy - Agoda consolidated multiple independent data pipelines into a central #ApacheSpark platform, eliminating financial data inconsistencies.
A multi-layered quality framework - with automated checks, ML anomaly detection, and data contracts - ensures accurate financial metrics while handling millions of daily bookings.
Deep dive into the architecture here ⇨ https://bit.ly/4a109NP
-
【ハンズオン】OCI Data Flowで始めるApache Spark|ETLからMLまで体験してみよう!
https://qiita.com/yushibats/items/559b65f72efeccf8865b?utm_campaign=popular_items&utm_medium=feed&utm_source=popular_items -
【ハンズオン】OCI Data Flowで始めるApache Spark|ETLからMLまで体験してみよう!
https://qiita.com/yushibats/items/559b65f72efeccf8865b?utm_campaign=popular_items&utm_medium=feed&utm_source=popular_items -
Apache Spark không tự động nhanh. Tốc độ của nó phụ thuộc vào cách dùng: tránh phân vùng sai, shuffle không cần thiết và lạm dụng cache. Hiểu rõ mô hình thực thi của Spark là chìa khóa để tối ưu hiệu suất.
#ApacheSpark #BigData #DataEngineering #Performance #Optimization #DuLieuLon #CongNgheDuLieu #HieuSuat #ToiUu
-
Apache Spark's new Declarative Pipelines framework simplifies ETL development by letting engineers focus on defining transformations while automating orchestration & error handling. This open-source solution handles batch/streaming workloads via Python/SQL interfaces, significantly reducing boilerplate code. Promising productivity gains for teams managing complex Spark pipelines. What impact might declarative approaches have on your workflow? #ApacheSpark #ETL #OpenSource
-
Discover how Decathlon, one of the world’s leading sports retailers, adopted the #opensource library #Polars to optimize its data workflows.
By migrating from Apache Spark to Polars for small input datasets, Decathlon achieved:
• Significant speed
• Meaningful cost savings👉 Learn more: https://bit.ly/4qmb2zc