use half of the available cores for rocBLAS - #281
Conversation
|
Let's give it a try bot: build repo:eessi.io-2025.06-software instance:eessi-bot-aws-eu-south on:arch=x86_64/amd/zen5 for:arch=x86_64/amd/zen5,accel=amd/gfx908+amd/gfx90a+amd/gfx942+amd/gfx1030+amd/gfx1100+amd/gfx1101+amd/gfx1200+amd/gfx1201 |
|
New job on instance
|
|
Not sure what's going on with these jobs. It doesn't look like they run out of memory, but it's getting stuck at some point, and ultimately Slurm aborts and requeues the job on the same node? On the node itself I only see these messages in the syslog: |
This should hopefully prevent it from running out of memory.