SemiAnalysis 对 GPU 集群健康检查做最基本测试:看它在纸面上是否可能跑通,结果发现好几个根本不可能成立。在 Amazon HyperPod Slurm 上首次测试时,一个健康检查要求节点先健康才能运行,但该检查本身又是节点重新入池所必需的,形成死循环。SemiAnalysis 称测试中约有五到十个这类"健康检查从未有机会运行"的例子,并质疑这些方案交付客户前是否经过充分验证。
The most basic test SemiAnalysis applies to a GPU cluster health check is whether it could possibly work on paper. Several could not.
"We've been through a few of these clusters where there's no way it could possibly work. These are health checks so bad they're worse than no health checks, because they're actively interfering with jobs."
"On Amazon HyperPod Slurm, the first time we tested, there was a health check that needed the node to be healthy before it could run. It could only run once the node was back in the fleet, and it was needed to bring the node back into the fleet."
"There are probably five or ten examples of this throughout testing where the health check never had a chance. It makes you wonder whether these people are really putting it through the paces before they hand it off to customers."
来源:SemiAnalysis · x.com