The failed component has been replaced, and the file system is now fully operational again.
The issue has been identified and mitigated. Full resolution will require a hardware part replacement, which will likely take place tomorrow.
In the meantime, the file system is up and running, and jobs should have resumed. We’ll monitor the situation until the hardware replacement is complete, and close this incident when it’s done.
A hardware-related failure occurred on the storage system that hosts $SCRATCH and $GROUP_SCRATCH. We’re investigating.