Pseudo-terminal will not be allocated because stdin is not a terminal. ******************************************************************************** * Welcome to * * _ _ ___ _______ _ ____ * * | | | | \ \ / / ____| | / ___| Juelich Wizard * * _ | | | | |\ \ /\ / /| _| | | \___ \ for * * | |_| | |_| | \ V V / | |___| |___ ___) | European Leadership * * \___/ \___/ \_/\_/ |_____|_____|____/ Science * * * ******************************************************************************** * Information about the system, latest changes, user documentation and FAQs: * * -> https://go.fzj.de/JUWELS -> https://go.fzj.de/juwels-known-issues * * JUWELS cluster slurm job reports: * * -> https://go.fzj.de/llview-juwels * ******************************************************************************** * Status information also at https://go.fzj.de/status-juwels-cluster * ******************************************************************************** 2026-06-12T13:00+0200 Information - Slurm instabilities This week there were 2 instances (one on Monday and one on Wednesday) where the slurm controller failed and stopped responding to commands. Friday after 13:00 the rate of these failures increased. The root cause is unfortunately not clear, but there are indications of a deadlock inside the controller code. No changes have been done on the system in the last weeks that could have contributed to this issue. This implies that until this is solved issuing slurm commands might ocassionally fail during some minutes, and that a system-wide reservation might be put in place if necessary to debug further. Running jobs are not affected. We have put measures in place to monitor the issue closer and have automatic recovery actions outside of working hours. -------------------------------------------------------------------------------- 2026-07-05T00:30+0200 Critical Incident - Cooling failure in the datacenter At 00:30 a general cooling failure on the datacenter resulted in an overheating of all systems installed in the 16.4 building. Cooling seemed restored at 03:30. The root cause is being investigated. As a result JUWELS Cluster is at the moment completely unavailable. JUWELS Booster and JURECA-DC are at the moment online, even though in the event JURECA- DC lost 2 complete racks (a CPU-only rack and a GPU rack). # Update 2026-07-05T08:15:00 JUWELS Booster and JURECA-DC are NOT online. The cooling seems stable but production won't be resumed until it is confirmed that the cooling infrastructure can handle the resulting load reliably. -------------------------------------------------------------------------------- Starting on 2026-07-07T09:00+0200 Planned - Software maintenance On 2026-07-07 there is a maintenance for software updates on JUWELS. Login nodes are updated, so a short interruption is to be expected. Compute nodes are updated and full availability is expected during the second half of the day. ******************************************************************************** +------------------------------------------------------------------------------+ | This node is in maintenance. Overall system state visible on the status page | +------------------------------------------------------------------------------+ Connection closed by 134.94.0.104 port 22