RWTH High Performance Computing (HPC)

You can find more information about the service in our documentation portal.

Nodes unavailable and Jobs stuck

Partial Outage
Mon 07/13/2026 12:50 PM - Unknown

We must report that compute nodes within Claix2023, Claix2025 as well as IH and private systems are currently suffering from random operating system failures that cause the jobs to become unavailable and for Jobs to hang while finishing in a completing state.
We know that this causes the waiting times to increase and that users are unable to stop their hanging jobs.
We are currently working on solving the problem with the uttermost priority.

Until the problem is resolved, we will manually stop users jobs in a stuck state and reset the nodes to a usable state manually.
This also takes time and work, so we ask users to please have patience until it is resolved.
To minimize node downtime, we ask users to try and use single node jobs where possible.

We apologize for this inconvenience and hope to have the issue resolved as soon as possible.

20.07.2026 13:02

ssh into Batchjobs disabled

Partial Outage
Thu 07/09/2026 08:00 AM - Unknown

Due to latent bugs in resource accounting, ssh-ing into nodes where one of your own jobs is running is temporarily disabled.

We are aware of the popularity of this feature and aim to reenable it in a future maintenance.

13.07.2026 13:44

MySQL Database Server Maintenance - MFA Restricted

Maintenance
Tue 07/28/2026 07:00 AM - Tue 07/28/2026 07:30 AM

During the specified maintenance window, the IT Center will perform maintenance on the MySQL databases used by the Multi-Factor Authentication (MFA) service. As part of the maintenance, the databases will be unavailable for a total of approximately 15 minutes.

During this time, new logins to services using RWTH Single Sign-On will not be possible. Existing sessions will remain active and will not be affected by the maintenance.

08.07.2026 13:00

copy23-2 temporarily unavailable

Partial Maintenance
Mon 06/29/2026 04:10 PM - Unknown

Due to maintenance work and debugging, copy23-2 is unavailable at the moment. Please use copy23-1 in the meantime.

29.06.2026 16:10

Recently expired reports

Login to CLAIX currently not possible - malfunction of frontend (dialog) nodes

Partial Outage
Mon 07/20/2026 01:00 PM - Tue 07/21/2026 11:57 AM

Due to a GPFS failure on the login nodes, it is currently not possible to log in, as the home directory cannot be loaded.

20.07.2026 14:00
Updates

Partial outage has been resolved.

21.07.2026 11:57

Access to CLAIX restricted due to Open Security Issues

Partial Outage
Wed 07/08/2026 02:15 PM - Tue 07/14/2026 12:54 PM

Due to unresolved security issues, the access to the CLAIX dialog systems is restricted at the moment. Updates are being prepared. Once deployed, the dialog nodes can be opened again.
Already submitted batch jobs are not affected and scheduled.

08.07.2026 14:24
Updates

We would like to inform users of the ongoing security updates we are deploying across the Claix Cluster.
These updates aim to patch a dangerous root exploit that could affect the entire system.
In addition to the existing login nodes, all compute nodes haven been drained to apply the necessary security patches.
Login nodes are patched and should be usable for users. The compute nodes on the other hand are being drained.
As individual compute nodes go idle and available for updates, we will deploy the security patches.
Since the progress is gradual as nodes go free and idle from the draining procedure, the time require to patch the entire system might take over a week.
If you are an HPC user with active jobs, we recommend you cancel any non essential jobs manually using scancel.
This will free the susceptible nodes sooner for security updates.

09.07.2026 16:18

up to now, 22 out of 48 GPU systems are still not updated as also 144 out of 661 HPC systems.
We are continually working on updating the systems as soon as they are empty.

10.07.2026 12:58

With the exception of a few systems, the problems have been resolved.

14.07.2026 12:54

[CLAIX-2025] Downtime due to CDU maintenance work; DGUV-Prüfung

Partial Maintenance
Tue 07/07/2026 09:00 AM - Thu 07/09/2026 10:42 AM

The CLAIX-2025 cluster must be shut-down due to maintenance work at the CDU.

06.07.2026 15:07
Updates

The "DGUV certification" measurements are also conducted; the maintenance needs to be extended for that purpose.

07.07.2026 11:03

The CDU's faulty pump was replaced. The DGUV measurements are still pending due to technical issues.

07.07.2026 16:03

Thde DGUV testing is now in progress.

08.07.2026 14:14

The DGUV tests are completed, the cluster will be powered-on shortly.

09.07.2026 10:43

[CLAIX-2025] Performance Tuning of HPC Fabric

Partial Maintenance
Tue 06/30/2026 01:25 PM - Mon 07/13/2026 01:00 PM

To resolve observed slow performance, all CLAIX-2025 nodes are under temporary maintenance during performance tuning and benchmarks. We will open the cluster again for the pilot users afterwards.

30.06.2026 13:24
Updates

The performance tuning is ongoing, but will be limited to allow for a pilot operation on CLAIX-2025.

13.07.2026 13:46