#slurm — Public Fediverse posts
Live and recent posts from across the Fediverse tagged #slurm, aggregated by home.social.
-
Ich bekomme noch die Krise mit #snakemake, #slurm und dem slurm-extra-plugin und #profiles.
Die Handles und Keys werden ständig unterschiedlich verwendet, je nach dem in welche Dokumentation man schaut. 😭
-
Ich bekomme noch die Krise mit #snakemake, #slurm und dem slurm-extra-plugin und #profiles.
Die Handles und Keys werden ständig unterschiedlich verwendet, je nach dem in welche Dokumentation man schaut. 😭
-
Ich bekomme noch die Krise mit #snakemake, #slurm und dem slurm-extra-plugin und #profiles.
Die Handles und Keys werden ständig unterschiedlich verwendet, je nach dem in welche Dokumentation man schaut. 😭
-
Ich bekomme noch die Krise mit #snakemake, #slurm und dem slurm-extra-plugin und #profiles.
Die Handles und Keys werden ständig unterschiedlich verwendet, je nach dem in welche Dokumentation man schaut. 😭
-
128-GPU LLM training stalled at 3 AM? 🛑 Escape the hardware telemetry trap.
100% GPU metrics lie during deadlocks. Our guide fixes it:
• Monitor power draw to expose collective deadlocks
• Catch thermal stragglers early via Prometheus
• Secure Ray dashboards using encrypted tunnelsDeploy on ServerMO GPU Dedicated Servers for unthrottled multi-node performance. ⚡
🔗 https://www.servermo.com/blogs/distributed-llm-training-slurm/
#Slurm #DevOps #SysAdmin #SRE #Linux #LLM #Telemetry #ServerMO
-
128-GPU LLM training stalled at 3 AM? 🛑 Escape the hardware telemetry trap.
100% GPU metrics lie during deadlocks. Our guide fixes it:
• Monitor power draw to expose collective deadlocks
• Catch thermal stragglers early via Prometheus
• Secure Ray dashboards using encrypted tunnelsDeploy on ServerMO GPU Dedicated Servers for unthrottled multi-node performance. ⚡
🔗 https://www.servermo.com/blogs/distributed-llm-training-slurm/
#Slurm #DevOps #SysAdmin #SRE #Linux #LLM #Telemetry #ServerMO
-
128-GPU LLM training stalled at 3 AM? 🛑 Escape the hardware telemetry trap.
100% GPU metrics lie during deadlocks. Our guide fixes it:
• Monitor power draw to expose collective deadlocks
• Catch thermal stragglers early via Prometheus
• Secure Ray dashboards using encrypted tunnelsDeploy on ServerMO GPU Dedicated Servers for unthrottled multi-node performance. ⚡
🔗 https://www.servermo.com/blogs/distributed-llm-training-slurm/
#Slurm #DevOps #SysAdmin #SRE #Linux #LLM #Telemetry #ServerMO
-
128-GPU LLM training stalled at 3 AM? 🛑 Escape the hardware telemetry trap.
100% GPU metrics lie during deadlocks. Our guide fixes it:
• Monitor power draw to expose collective deadlocks
• Catch thermal stragglers early via Prometheus
• Secure Ray dashboards using encrypted tunnelsDeploy on ServerMO GPU Dedicated Servers for unthrottled multi-node performance. ⚡
🔗 https://www.servermo.com/blogs/distributed-llm-training-slurm/
#Slurm #DevOps #SysAdmin #SRE #Linux #LLM #Telemetry #ServerMO
-
I have forked "Slurm in Kubernetes". As far as I can tell it was abandoned. I made some updates and got it working in minikube: https://codeberg.org/invirtuate/slik
-
I have forked "Slurm in Kubernetes". As far as I can tell it was abandoned. I made some updates and got it working in minikube: https://codeberg.org/invirtuate/slik
-
I like that "lazy version" exists for many things! 😍
🌀 **lazyslurm** — A TUI for monitoring and managing Slurm clusters
💯 Track jobs, inspect nodes, tail logs, monitor partitions, & manage workloads
🦀 Written in Rust & built with @ratatui_rs
⭐ GitHub: https://github.com/hill/lazyslurm
-
I like that "lazy version" exists for many things! 😍
🌀 **lazyslurm** — A TUI for monitoring and managing Slurm clusters
💯 Track jobs, inspect nodes, tail logs, monitor partitions, & manage workloads
🦀 Written in Rust & built with @ratatui_rs
⭐ GitHub: https://github.com/hill/lazyslurm
-
It is quite fun to occasionally come back to using #slurm to send commands to a super computer node.
Let's see if running a #TopicModelling script with a V100 GPU reduces the running time from 80+ hours to a few minutes, as I expect :)
-
It is quite fun to occasionally come back to using #slurm to send commands to a super computer node.
Let's see if running a #TopicModelling script with a V100 GPU reduces the running time from 80+ hours to a few minutes, as I expect :)
-
Scaling Compute: The Friction of Distributed AI Infrastructure
On 23 May 2026, AI companies face hardware delays. Learn why complex multi-node GPU clusters are failing and how this impacts industrial AI production.
#aiinfrastructure, #gpuclusters, #slurm, #industrialai, #techupdate
https://newsletter.tf/ai-cluster-hardware-failures-may-2026/
-
Industrial AI clusters are struggling with stability today. This is a major change from last year when software speed was the only focus.
#aiinfrastructure, #gpuclusters, #slurm, #industrialai, #techupdate
https://newsletter.tf/ai-cluster-hardware-failures-may-2026/ -
Today I've done:
- accepted a review request (difficult to even see them these days in the wave of other mails). It is from some colleagues at a place where I have some acquaintances. None of the persons I know, so I am lucky for otherwise I would have declined. As the review is anonymous, and I know many people in the field, mentioning this here on Mastodon will not become an issue.
- followed an online meeting. Actually one of the non-boring ones.
- re-installed an environment I involuntarily screwed up yesterday when testing things for a user
- had a nice lunch with @KrawallHamster , @moschlar and others
- written a number of mails
- debugged code, debugged code, debugged code - will postpone a new #Snakemake plugin release for #SLURM to next week, when I can think straight. At least the CI pipeline is fine for this PR I was working on. But I always do live tests on an actual cluster if the change is not trivial.
- actually finished the review task (first round)And I 🚴 , up- and downhill, through the May heat (this is a thing these days!!!). Lesson learned: Next time, I will take a break, sit on a bench, read and drink to have a rest for the last leg. The afternoon heat is no fun!
-
Today I've done:
- accepted a review request (difficult to even see them these days in the wave of other mails). It is from some colleagues at a place where I have some acquaintances. None of the persons I know, so I am lucky for otherwise I would have declined. As the review is anonymous, and I know many people in the field, mentioning this here on Mastodon will not become an issue.
- followed an online meeting. Actually one of the non-boring ones.
- re-installed an environment I involuntarily screwed up yesterday when testing things for a user
- had a nice lunch with @KrawallHamster , @moschlar and others
- written a number of mails
- debugged code, debugged code, debugged code - will postpone a new #Snakemake plugin release for #SLURM to next week, when I can think straight. At least the CI pipeline is fine for this PR I was working on. But I always do live tests on an actual cluster if the change is not trivial.
- actually finished the review task (first round)And I 🚴 , up- and downhill, through the May heat (this is a thing these days!!!). Lesson learned: Next time, I will take a break, sit on a bench, read and drink to have a rest for the last leg. The afternoon heat is no fun!
-
Today:
- written a blurb for my presentation at the upcoming #nanopub session
- release the #SLURM executor plugin for #Snakemake v2.7.0 - see https://fediscience.org/@snakemake/116617420491776431
- tried to mitigate the issue that TMOUT on an HPC login brings: sending SIGHUB to all detached multiplexers (so far no remedy and I tried a lot(!), don't send me tips).
- futile further debugging attempts. In the end it worked. Might result in a new release next week. -
Today:
- written a blurb for my presentation at the upcoming #nanopub session
- release the #SLURM executor plugin for #Snakemake v2.7.0 - see https://fediscience.org/@snakemake/116617420491776431
- tried to mitigate the issue that TMOUT on an HPC login brings: sending SIGHUB to all detached multiplexers (so far no remedy and I tried a lot(!), don't send me tips).
- futile further debugging attempts. In the end it worked. Might result in a new release next week. -
От майнинга на попутном газе к AI-фабрикам: история Crusoe
У AI-индустрии есть серьезная проблема: как развернуть вычислительную инфраструктуру раньше и быстрее (да еще и дешевле) конкурентов? Основной дефицитный ресурс сейчас — электричество, а не чипы или их компоненты, как вы могли предположить. Техногиганты думают, где поставить стойки, чем их охлаждать, но главное, где взять энергию, чтобы питать всю AI-систему. И у одного стартапа из Денвера есть нестандартное решение — портативные модульные AI-дата-центры, которые можно размещать в самых нестандартных условиях. Компания пришла в ИТ из мира крипты: изначально она вела деятельность установкой майнинг-машин, которые брали энергию от попутного газа на нефтяных вышках. Сегодня я расскажу вам о компании Crusoe — которая крайне нестандартно превращает энергию в вычислительную мощность. Разберем их бизнес-модель и поймем, что такое вертикально интегрированная AI-инфраструктура.
https://habr.com/ru/companies/ru_mts/articles/1022116/
#Crusoe #AIинфраструктура #датацентры #GPUоблако #облачные_вычисления #inference #Kubernetes #Slurm #edge_computing #энергетика
-
От майнинга на попутном газе к AI-фабрикам: история Crusoe
У AI-индустрии есть серьезная проблема: как развернуть вычислительную инфраструктуру раньше и быстрее (да еще и дешевле) конкурентов? Основной дефицитный ресурс сейчас — электричество, а не чипы или их компоненты, как вы могли предположить. Техногиганты думают, где поставить стойки, чем их охлаждать, но главное, где взять энергию, чтобы питать всю AI-систему. И у одного стартапа из Денвера есть нестандартное решение — портативные модульные AI-дата-центры, которые можно размещать в самых нестандартных условиях. Компания пришла в ИТ из мира крипты: изначально она вела деятельность установкой майнинг-машин, которые брали энергию от попутного газа на нефтяных вышках. Сегодня я расскажу вам о компании Crusoe — которая крайне нестандартно превращает энергию в вычислительную мощность. Разберем их бизнес-модель и поймем, что такое вертикально интегрированная AI-инфраструктура.
https://habr.com/ru/companies/ru_mts/articles/1022116/
#Crusoe #AIинфраструктура #датацентры #GPUоблако #облачные_вычисления #inference #Kubernetes #Slurm #edge_computing #энергетика
-
От майнинга на попутном газе к AI-фабрикам: история Crusoe
У AI-индустрии есть серьезная проблема: как развернуть вычислительную инфраструктуру раньше и быстрее (да еще и дешевле) конкурентов? Основной дефицитный ресурс сейчас — электричество, а не чипы или их компоненты, как вы могли предположить. Техногиганты думают, где поставить стойки, чем их охлаждать, но главное, где взять энергию, чтобы питать всю AI-систему. И у одного стартапа из Денвера есть нестандартное решение — портативные модульные AI-дата-центры, которые можно размещать в самых нестандартных условиях. Компания пришла в ИТ из мира крипты: изначально она вела деятельность установкой майнинг-машин, которые брали энергию от попутного газа на нефтяных вышках. Сегодня я расскажу вам о компании Crusoe — которая крайне нестандартно превращает энергию в вычислительную мощность. Разберем их бизнес-модель и поймем, что такое вертикально интегрированная AI-инфраструктура.
https://habr.com/ru/companies/ru_mts/articles/1022116/
#Crusoe #AIинфраструктура #датацентры #GPUоблако #облачные_вычисления #inference #Kubernetes #Slurm #edge_computing #энергетика
-
RE: https://fediscience.org/@snakemake/116295568336688286
This is a big step forward: The SLURM plugin for Snakemake now supports so-called job arrays. These are cluster jobs, with ~ equal resource requirements in terms of memory and compute resources.
The change in itself was big: The purpose of a workflow system is to make use of the vast resources of an HPC cluster. Hence, jobs are submitted to run concurrently. However, for a job array, we have to "wait" for all eligible jobs to be ready. And then we submit.
To preserve concurrent execution of other jobs which are ready to be executed, a thread pool has been introduced. In itself, I do not see job arrays as such a big feature: The LSF system profited much more from arrays than the rather lean SLURM implementation does.
BUT: the new code base will ease further development to pooling many shared memory tasks (applications which support no parallel execution or are confined to one computer by "only" supporting threading). Until then, there is more work to do.
#HPC #SLURM #Snakemake #SnakemakeHackathon2026 #ReproducibleComputing #OpenScience
-
RE: https://fediscience.org/@snakemake/116295568336688286
This is a big step forward: The SLURM plugin for Snakemake now supports so-called job arrays. These are cluster jobs, with ~ equal resource requirements in terms of memory and compute resources.
The change in itself was big: The purpose of a workflow system is to make use of the vast resources of an HPC cluster. Hence, jobs are submitted to run concurrently. However, for a job array, we have to "wait" for all eligible jobs to be ready. And then we submit.
To preserve concurrent execution of other jobs which are ready to be executed, a thread pool has been introduced. In itself, I do not see job arrays as such a big feature: The LSF system profited much more from arrays than the rather lean SLURM implementation does.
BUT: the new code base will ease further development to pooling many shared memory tasks (applications which support no parallel execution or are confined to one computer by "only" supporting threading). Until then, there is more work to do.
#HPC #SLURM #Snakemake #SnakemakeHackathon2026 #ReproducibleComputing #OpenScience
-
A few #slurm tidbits:
Total submitted jobs per user, sorted:
```
squeue | sed 's/ \+/\t/g' | cut -f5 \
| sort | uniq -c | sort -hr
```Running jobs per user:
```
squeue | grep ' R ' | sed 's/ \+/\t/g' \
| cut -f5 | sort | uniq -c | sort -hr
```Pending jobs per user:
```
squeue | grep ' PD ' | sed 's/ \+/\t/g' \
| cut -f5 | sort | uniq -c | sort -hr
``` -
A few #slurm tidbits:
Total submitted jobs per user, sorted:
```
squeue | sed 's/ \+/\t/g' | cut -f5 \
| sort | uniq -c | sort -hr
```Running jobs per user:
```
squeue | grep ' R ' | sed 's/ \+/\t/g' \
| cut -f5 | sort | uniq -c | sort -hr
```Pending jobs per user:
```
squeue | grep ' PD ' | sed 's/ \+/\t/g' \
| cut -f5 | sort | uniq -c | sort -hr
``` -
As for the little executor plugin for the #SLURM batch system (for which I promised a release supporting array job support) ... Well, only a little bug fix release could be accomplished: https://github.com/snakemake/snakemake-executor-plugin-slurm/releases/tag/v2.5.4
Unfortunately, I wanted to use the common #Snakemake logo without the letters "#HPC" and missed one entry. So our announcement bot did not work.
Anyway, a faulty file system connection kept me from debugging the new feature. Stay tuned. It is almost ready.
-
As for the little executor plugin for the #SLURM batch system (for which I promised a release supporting array job support) ... Well, only a little bug fix release could be accomplished: https://github.com/snakemake/snakemake-executor-plugin-slurm/releases/tag/v2.5.4
Unfortunately, I wanted to use the common #Snakemake logo without the letters "#HPC" and missed one entry. So our announcement bot did not work.
Anyway, a faulty file system connection kept me from debugging the new feature. Stay tuned. It is almost ready.
-
Finally, some personal progress: Thanks to @fbartusch a bug of the #SLURM executor plugin for Snakemake was fixed (dealing with nested quoting). A release is upcoming.
And: I generated my first (still faulty) test #nanopub from Snakemake 🥳
-
Finally, some personal progress: Thanks to @fbartusch a bug of the #SLURM executor plugin for Snakemake was fixed (dealing with nested quoting). A release is upcoming.
And: I generated my first (still faulty) test #nanopub from Snakemake 🥳
-
This cannot be:
I am trying to compile a few stats for the #Snakemake executor plugin for #SLURM on #HPC systems. Preparing for a lighting talk at the #SnakemakeHackathon2026
PyPi: 20,000 downloads last month
BioConda: > 60,000 total (aggregated over all versions)Impressive as it might be, this is contradictory. PyPi would exceed BioConda by a huge margin.
Does anyone know how to get all-time statistics from either platform? #BioConda or #PyPi?
-
This cannot be:
I am trying to compile a few stats for the #Snakemake executor plugin for #SLURM on #HPC systems. Preparing for a lighting talk at the #SnakemakeHackathon2026
PyPi: 20,000 downloads last month
BioConda: > 60,000 total (aggregated over all versions)Impressive as it might be, this is contradictory. PyPi would exceed BioConda by a huge margin.
Does anyone know how to get all-time statistics from either platform? #BioConda or #PyPi?
-
The #Snakemake plugin for #SLURM on #HPC clusters will support JobArrays, soon:
1057691_1 2dcf44cc-+ rule_map_reads_wild+ 32 COMPLETED 0:0
1057691_2 2dcf44cc-+ 32 RUNNING 0:0
1057691_3 2dcf44cc-+ 32 RUNNING 0:0
1057691_4 2dcf44cc-+ 32 RUNNING 0:0
1057691_5 2dcf44cc-+ 32 RUNNING 0:0
1057691_6 2dcf44cc-+ 32 RUNNING 0:0Hope to do more during next week's #SnakemakeHackathon2026 / #SnakemakeHackathon
-
The #Snakemake plugin for #SLURM on #HPC clusters will support JobArrays, soon:
1057691_1 2dcf44cc-+ rule_map_reads_wild+ 32 COMPLETED 0:0
1057691_2 2dcf44cc-+ 32 RUNNING 0:0
1057691_3 2dcf44cc-+ 32 RUNNING 0:0
1057691_4 2dcf44cc-+ 32 RUNNING 0:0
1057691_5 2dcf44cc-+ 32 RUNNING 0:0
1057691_6 2dcf44cc-+ 32 RUNNING 0:0Hope to do more during next week's #SnakemakeHackathon2026 / #SnakemakeHackathon
-
I had the chance to present a poster about our HPC cluster BinAC 2 at #deRSE26
We're providing computational resources for researchers in Baden-Württemberg working in the fields of Bioinformatics, Astrophysics, Geosciences, Pharmacy and Medical Informatics:
https://wiki.bwhpc.de/e/BinAC2If you're a researcher at an university in Baden-Württemberg from another field, check out the other HPC clusters bwHPC provides:
https://www.bwhpc.de/cluster.phpPoster on Zenodo:
https://zenodo.org/records/18860391#hpc #bioinformatics #astrophysics #geosciences #bwhpc #lustre #Slurm
-
On a similar note: there is another (draft) PR. The #SLURM executor plugin for #Snakemake is capable of respecting partition definitions since v. 2.
I had the notion, that this is rather difficult to set this up manually and wrote a little command line helper. It queries the SLURM config and writes out a preliminary partition configuration template. This still requires manual adaptation, I'm afraid.
A small step forward as it requires both an understanding of Snakemake and your local SLURM setup. The world is as is it is, the phantasy of admin teams is unlimited and a one-fits-all solution is not on the horizon.
Still, if you want to try it out and provide feedback, this would be very much appreciated! All suggestions are welcome!
-
On a similar note: there is another (draft) PR. The #SLURM executor plugin for #Snakemake is capable of respecting partition definitions since v. 2.
I had the notion, that this is rather difficult to set this up manually and wrote a little command line helper. It queries the SLURM config and writes out a preliminary partition configuration template. This still requires manual adaptation, I'm afraid.
A small step forward as it requires both an understanding of Snakemake and your local SLURM setup. The world is as is it is, the phantasy of admin teams is unlimited and a one-fits-all solution is not on the horizon.
Still, if you want to try it out and provide feedback, this would be very much appreciated! All suggestions are welcome!
-
I want to reach out: I have this pending release for the SLURM executor (https://github.com/snakemake/snakemake-executor-plugin-slurm/pull/412 ). It implements better error feedback (in case of hardware failures and otherwise). It would need some thorough checking, and I cannot provoke too many hardware failures. Is anyone working on an older cluster? 😉
Feedback (as issues) would be appreciated. Also nice, if you tell me it is working, here.
-
I want to reach out: I have this pending release for the SLURM executor (https://github.com/snakemake/snakemake-executor-plugin-slurm/pull/412 ). It implements better error feedback (in case of hardware failures and otherwise). It would need some thorough checking, and I cannot provoke too many hardware failures. Is anyone working on an older cluster? 😉
Feedback (as issues) would be appreciated. Also nice, if you tell me it is working, here.
-
PSA for my #HPC cluster operators out there. A new CVE was announced for #MUNGE, a popular authentication mechanism used in #Slurm
https://github.com/dun/munge/security/advisories/GHSA-r9cr-jf4v-75gh
-
PSA for my #HPC cluster operators out there. A new CVE was announced for #MUNGE, a popular authentication mechanism used in #Slurm
https://github.com/dun/munge/security/advisories/GHSA-r9cr-jf4v-75gh
-
What Does #Nvidia’s Acquisition of #SchedMD Mean for #Slurm?
Slurm was developed at LLNL in the early 2000s to replace commercial workload management software for #HPC clusters and #supercomputers.
Addison Snell, the CEO of Intersect360, says the acquisition of SchedMD makes sense considering the emerging focus on developing #AI models to accelerate scientific discovery and engineering, and the need to integrate traditional HPC workloads and new AI ones.
https://www.hpcwire.com/2026/01/06/what-does-nvidias-acquisition-of-schedmd-mean-for-slurm/ -
What Does #Nvidia’s Acquisition of #SchedMD Mean for #Slurm?
Slurm was developed at LLNL in the early 2000s to replace commercial workload management software for #HPC clusters and #supercomputers.
Addison Snell, the CEO of Intersect360, says the acquisition of SchedMD makes sense considering the emerging focus on developing #AI models to accelerate scientific discovery and engineering, and the need to integrate traditional HPC workloads and new AI ones.
https://www.hpcwire.com/2026/01/06/what-does-nvidias-acquisition-of-schedmd-mean-for-slurm/ -
Does anyone here use the #Slurm `nss_slurm` extension?
I see in Slurm's documentation an example of how to enable the extension, but I can't find any examples of the referenced /etc/nss_slurm.conf file anywhere...
The source code of the extension seems to indicate that it is very simple file - e.g.
```
NodeName=<nodename>
SlurmdSpoolDir=<dir>
```but I just want an example to ensure that my assumptions are correct 😅
-
Does anyone here use the #Slurm `nss_slurm` extension?
I see in Slurm's documentation an example of how to enable the extension, but I can't find any examples of the referenced /etc/nss_slurm.conf file anywhere...
The source code of the extension seems to indicate that it is very simple file - e.g.
```
NodeName=<nodename>
SlurmdSpoolDir=<dir>
```but I just want an example to ensure that my assumptions are correct 😅
-
Nvidia’s acquisition of SchedMD, the company behind Slurm, is a strategic move that goes far beyond GPUs.
Slurm (Simple Linux Utility for Resource Management) is the de facto open-source workload manager for large-scale GPU clusters, widely used in supercomputing centers, AI labs, hyperscalers, and cloud GPU operators. It plays a critical role in ...
#NVIDIA #AIInfrastructure #OpenSource #Slurm #HPC #GPUs #AITraining #CloudComputing #tech #DataCenters
-
Nvidia’s acquisition of SchedMD, the company behind Slurm, is a strategic move that goes far beyond GPUs.
Slurm (Simple Linux Utility for Resource Management) is the de facto open-source workload manager for large-scale GPU clusters, widely used in supercomputing centers, AI labs, hyperscalers, and cloud GPU operators. It plays a critical role in ...
#NVIDIA #AIInfrastructure #OpenSource #Slurm #HPC #GPUs #AITraining #CloudComputing #tech #DataCenters