怎样看待美国最新超算 Summit( 四 )
我这里应该提到我的同事Jack Dongarra写过太湖之光不是基于Alpha处理器的。说实话,他应该问问他手下的大陆研究生去读一下那些没被翻成英语的太湖之光的文章。中文版是有说用了Alpha的,但是英文翻译就没了。我周围的懂双语的新加坡人确认了此事!设计师可能是害怕会被人指责说是偷窃了美国技术,所以就没有提到设计的这方面。他们这么做很不应该。太湖之光的创新足够让设计师拿到不止一个,两个Gordon Bell奖。
Each processor of Taihu Light looks like the Cray T3D on a chip. The Cray T3D was a nimble system based on Alpha processors that many HPC people feel was one of the best-designed supercomputers of all time.
太湖之光的每个处理器都像是一个芯片上集成了一整台Cray T3D。Cray T3D是一个很牛逼的机器,基于Alpha处理器,很多超算人认为这是有史以来设计的最好的超算之一。
Most supercomputers are severely communication-bound; the T3D was much less so, with an unusually good system balance and low-latency interconnect that made it easier to sustain a high fraction of the peak rated speed. Imagine a 256-processor T3D on a single chip (together with four processors that service that array), and a cleverly-cooled cabinet that packs hundreds of those close together, and a roomful of those cabinets, and you have a system that makes the DOE and NASA labs in the USA go… *gulp*.
绝大多数的超算严重的受到通讯的限制。T3D很大程度上克服了这点,他们有一个非常好的系统平衡,以及低延迟的互联系统,这让它能够很容易的维持在顶峰速度的高荷载上。想象一个256核的T3D集成在单片上(同时还有四个处理器协同阵列),以及一个设计的很聪明的冷却机柜,将几百片这样的处理器塞在一起,再把一个房间装满这样的机柜,你就有了一台能够让能源部和NASA实验室羡慕嫉妒恨的机器。
If the USA wants to really get back in the game and not just play catch-up, they need to break the me-too paradigm of filling standard racks with x86 processors that have GPU accelerators attached. We can get an order of magnitude improvement in operations per joule by rethinking everything. If I were doing it, I’d use a Very-Long Instruction Word (VLIW) processor with no caches, no instruction lookahead or speculative execution or branch prediction, explore the use of gallium nitride with 16-level logic instead of silicon CMOS, change the numeric representation from IEEE floats to posit arithmetic, connect the cabinets with free-space optics at terabytes per second per channel and a full crossbar, use only stacked memory and extensive use of in-processor RAM and ROM, and declare it a “moon shot” to make such a system work by 2022. This is the way all the great breakthroughs in supercomputing have been made historically… by being willing to change paradigms. Right now, the Chinese are proving better at breaking from legacy thinking than the USA.
如果美国真的想要胜利,而不是追赶,他们就不应该局限于那种“我也行”的思路,把一大堆带GPU加速器的x86处理器塞在一起。如果我们重新思考,就能够将每瓦的计算力提高一个数量级。如果我来干这事情的话,我会用一个超长指令处理器(VLIW),不带缓存,不要向前检测,推测执行或者分支预测,探索使用16层逻辑的GaN而不是硅CMOS,将数字表示从IEEE浮点换成假定算式,将机柜用free-space光纤连接起来,带宽在TB级别,只用堆叠内存和片上RAM和ROM,宣布这是个“革命性产品”并且节点在2022年运行。(译者:我不懂超算所以这段翻得不好大家姑且看看)这是历史上所有超算巨大突破的路径——愿意去做范式更新。目前,中国人证明他们比美国人更有勇气革新。
推荐阅读
- 聪明人养花,这3种“花”怎样也要养一盆,每年能省不少医药费
- 互联网怎样解决“家政服务上门速度慢”的问题
- 怎样看待从1月8号起,QQ钱包开始提现收费
- 银行it人怎样转型
- 德洛斯科夫雷斯|
- 中东问题|
- 汽车|冬天怎样让车内温度快速升高?座椅加热的最佳使用方式二,外循环的作用总结
- 怎样进入通信行业
- 怎样评价扶他柠檬茶的小说《云养汉》的结尾
- 骨裂|
