当前位置:首页 > 报告详情

MACO:用于 DNN 加速器的 HW-Mapping 协同优化框架.pdf

上传人: 芦苇 编号:651777 2025-05-01 22页 1.08MB

1、MACO:A HW-Mapping Co-optimization Framework for DNN AcceleratorsSpeaker:Wujie Zhong1The Hong Kong University of Science and Technology(Guangzhou)Guangzhou,ChinaCatalogue Introduction Related Works MACO Experiment Conclusion2Introduction DNN accelerators3GPU Tensor CoreTPUIntroduction Design Space Ex

2、ploration The capacity of data buffers The number of PEs The number of MACs Loop boundaries Loop order Tradeoff between power,performance and area(PPA)4Hardware SpaceMappingSpaceIntroduction Hardware Space Exploration5The architecture of a CNN accelerator:Simba MICRO19Introduction Hardware Space Exp

3、loration Explore the hardware parameters Computation bound More PEs or more MACs Memory bound Higher bandwidth and larger buffers6Introduction Mapping Space Exploration Loop nest7Introduction Mapping Space Exploration Six Memory Levels L0:PE Weight Register Level L1:PE Accumulator Buffer Level L2:PE

4、 Weight Buffer Level L3:PE Input Buffer Level L4:Global Buffer Level L5:DRAM Level8Introduction Mapping Space Exploration9A part of an example about mapping a convolution layer into a Simba-like chipletIntroduction Mapping Space Exploration Buffer Capacity Constraint10PE Weight Buffer:2 2 2 2 Relate

5、d Works Mapping Space Exploration Timeloop ISPASS19:exhaustive and random search Challenge:huge design space GAMMA ICCAD20:genetic algorithm CoSA ISCA21:Mixed Integer Programming(MIP)LEMON CF23:Mixed Integer Programming(MIP)11Related Works Hardware-Mapping Co-optimization DiGamma DATE22:genetic algo

6、rithm Target on a two-level memory hardware DOSA MICRO23:gradient-based methods Target on a single objective MEDEA DATE22:genetic algorithm Suboptimal solutions12MACO Overview13MACO Hardware Space Search Block Multi-objective Bayesian optimization(MOBO)Evaluat

word格式文档无特别注明外均可编辑修改,预览文件经过压缩,下载原文更清晰!
三个皮匠报告文库所有资源均是客户上传分享,仅供网友学习交流,未经上传用户书面授权,请勿作商用。
本文介绍了MACO框架,一种针对深度神经网络加速器的硬件映射共优化框架。主要内容包括:1)深度神经网络加速器的硬件空间探索,如Simba架构等;2)设计空间探索,包括数据缓冲区容量、PE数量、MAC数量、循环边界和循环顺序等的权衡;3)硬件映射共优化,包括遗传算法、混合整数规划等方法。实验部分设置了DNN模型、模拟器和搜索算法,结果表明,MACO在速度、能效和延迟等方面具有显著优势。
"MACO框架如何优化DNN加速器?" "如何通过MACO框架实现硬件与映射空间的协同优化?" "MACO框架在实验中取得了哪些显著成果?"
客服
商务合作
小程序
服务号
折叠