Files
matmul-analysis/BMM/BMM_Theory/examples/result_recommend.csv

8.1 KiB

1case_idbatch_abatch_bmnkdtype_adtype_bdtype_ctrans_atrans_bhas_biasout_nddeterministic_levelplan_case_idplan_branchplan_npuplan_opplan_used_core_numplan_split_bplan_m_cntplan_n_cntplan_grid_kplan_core_mapplan_b_coreplan_merge_b0plan_single_core_mplan_single_core_nplan_single_core_kplan_k_l1plan_b_l1plan_l1_formplan_base_mplan_base_nplan_base_kplan_l2_policy_inplan_l2_policy_outplan_swizzle_wplan_workspace_bytesplan_tail_strategyplan_tail_m_cntplan_tail_n_cntplan_tail_k_cntplan_tail_m_mainplan_tail_n_mainplan_tail_block_cntplan_tail_wave_numplan_fixpipe_unitflagplan_out_dtype_bytesplan_notegm_read_bytesl2_read_bytest_mte2_gmt_mte2_l2t_mte2dma_cmd_countt_dma_cmdcube_flopst_mmadfixpipe_bytest_fixpipet_reducet_steadyt_draint_totalbottleneckfeasibleviolationsbound_typeadvice
2to_matmul_demo11204820482048bf16bf16bf16FalseFalseFalseTrue0to_matmul_demo转MatmulAscend950PRbatch_mat_mul_v3321001折叠为 Matmul [2048,2048]x[2048,2048], 复用 Matmul 切分体系010020480100000转Matmul后由 Matmul 体系决定1110000True2BatchB=1免费折叠: 左矩阵 [1,2048,2048] 视图折叠为 [2048,2048], 零重排零 split167772160.01.048576e-050.01.048576e-050.00.017179869184.03.534952506995885e-0583886085.24288e-060.03.534952506995885e-050.03.534952506995885e-05MMADTrue计算BoundBatchA=1或BatchB=1, 折叠转普通Matmul
3special_k0_demo1281282562560bf16bf16bf16FalseFalseFalseTrue0special_k0_demo特殊分支Ascend950PRbatch_mat_mul_v3641111AIV 核间按行均分 (无 Cube tile 概念)0100001UB驻留(AIV)000allocatedirect_gm00不涉及(AIV逐元素)1110000False2K=0纯写值: 无任何计算, C=bias 或 0, 纯 AIV 写值; 按行均分到 AIV 核0.00.00.00.00.00.00.00.00.0167772161.048576e-050.01.048576e-050.01.048576e-05FIXPIPETrue写出BoundK=0纯写值
4special_k1_demo1281282562561bf16bf16bf16FalseFalseFalseTrue0special_k1_demo特殊分支Ascend950PRbatch_mat_mul_v3641111AIV 核间按行均分 (无 Cube tile 概念)0100101UB驻留(AIV) UB乒乓000allocatedirect_gm00不涉及(AIV逐元素)1110000False2K=1逐元素乘: 退化为 C=A⊙B 无累加深度, Cube 16x16x16 粒度浪费 15/16; 走 AIV 通路 GM->UB->Mul->GM, UB乒乓 (B>=2*AIV 双batch乒乓流水)1310720.08.192e-080.08.192e-080.00.08388608.06.206060606060606e-07167772161.048576e-050.01.048576e-050.01.048576e-05FIXPIPETrue写出BoundK=1逐元素乘, 走AIV向量通路
5merge_demo_k_trunc204820483232256bf16bf16bf16FalseFalseFalseTrue0merge_demo_k_truncMergeBatchAscend950PRbatch_mat_mul_v33232111切B均分(核间零重复读零依赖)6441281282562568合并驻留128128128allocate(GM->L1随路驻留L2)direct_gm00不涉及(核内不切M/N)1110000True2b0=4 (L0C上限5.7/算存比上限19.0/b_core=64); K截断; 合并后单次DMA搬入 A'[128,256]+B'[256,128]20971520.04.194304e-050.04.2743039999999997e-0516.08.000000000000001e-07134217728.08.837381267489712e-061310722.62144e-060.04.2743039999999997e-053.0192408230452674e-074.3044964082304525e-05MTE2_GMTrue访存Bound(GM)两分支均合法, 仲裁: [分界条件] MergeBatch最优=True (k_L1=K(截断); b_core=64 vs 阈值 b0*(T_comp+T_write)/T_cmd=6.0; drain惩罚=(b0-1)*(T_comp+T_write)=0.23us, 搬移节省=b_core*(1-1/b0)*T_cmd=2.40us); [时延模型] T_MergeBatch=43.04us vs T_IterBatch=45.22us -> MergeBatch更优; [裁决] MergeBatch
6merge_iter_arbitrate1281286464512bf16bf16bf16FalseFalseFalseTrue0merge_iter_arbitrateIterBatchAscend950PRbatch_mat_mul_v33232111切B轮转分配(核间零重复读零依赖)4164645125122b_双batch乒乓6464256allocate(GM->L1随路驻留L2)direct_gm(输出仅写一次,直写GM不占L2)00不涉及(核内不切M/N)1110000True2双batch乒乓: 2*(MK+KN)*dtype=256KB <= L15242880.01.048576e-050.01.0685759999999999e-0542.0000000000000002e-0716777216.01.104672658436214e-06327686.5536e-070.01.0685759999999999e-054.400081646090535e-071.1125768164609053e-05MTE2_GMTrue访存Bound(GM)两分支均合法, 仲裁: [分界条件] MergeBatch最优=False (k_L1=K(截断); b_core=4 vs 阈值 b0*(T_comp+T_write)/T_cmd=17.6; drain惩罚=(b0-1)*(T_comp+T_write)=0.44us, 搬移节省=b_core*(1-1/b0)*T_cmd=0.10us); [时延模型] T_MergeBatch=11.49us vs T_IterBatch=11.13us -> IterBatch更优; [裁决] IterBatch
7iter_demo_form_b1281286464256bf16bf16bf16FalseFalseFalseTrue0iter_demo_form_bIterBatchAscend950PRbatch_mat_mul_v33232111切B轮转分配(核间零重复读零依赖)4164642562562b_双batch乒乓6464256allocate(GM->L1随路驻留L2)direct_gm(输出仅写一次,直写GM不占L2)00不涉及(核内不切M/N)1110000True2双batch乒乓: 2*(MK+KN)*dtype=128KB <= L12621440.05.24288e-060.05.4428799999999995e-0642.0000000000000002e-078388608.05.52336329218107e-07327686.5536e-070.05.4428799999999995e-063.0192408230452674e-075.744804082304526e-06MTE2_GMTrue访存Bound(GM)仅 IterBatch 条件满足
8iter_demo_form_d646464648192bf16bf16bf16FalseFalseFalseTrue0iter_demo_form_dIterBatchAscend950PRbatch_mat_mul_v33232111切B轮转分配(核间零重复读零依赖)216464819210241d_两侧都切K6464256allocate(GM->L1随路驻留L2)direct_gm(输出仅写一次,直写GM不占L2)00不涉及(核内不切M/N)1110000True2两侧都切K: k_L1=1024, K段成对流水, batch边界天然无缝; dValueA=2048B/dValueB=128B41943040.08.388608e-050.08.468608e-05168.000000000000001e-07134217728.08.837381267489712e-06163843.2768e-070.08.468608e-057.16176329218107e-078.540225632921811e-05MTE2_GMTrue访存Bound(GM)仅 IterBatch 条件满足
9streamk_demo4412812810240bf16bf16bf16FalseFalseFalseTrue0streamk_demoStreamKAscend950PRbatch_mat_mul_v33211132B/M/N切出4块, 每块32核切K归约 (归约组内核c负责K段[c*K/32,(c+1)*K/32))111281283202561K段标准分块流水12812864allocate(部分和驻留L2)resident(部分和4B驻留L2, 防精度丢失不随C的fp16/fp8转换)08388608grid_K=32路切K+归约11320000True4P=1.00, grid_K=32, 部分和驻留L2按4B写出, AIV归约后按C dtype=2B写最终327680.00.06.5536e-060.06.5536e-060.00.041943040.02.761681646090535e-060.00.03.4067453613053613e-066.5536e-063.4067453613053613e-069.960345361305361e-06MTE2_GMTrue访存Bound(GM)P<=C/2, B/M/N并行度买不满, 切K (grid_K=32)
10asw_demo_full22819281921024bf16bf16bf16FalseFalseFalseTrue0asw_demo_fullASW_BasicAscend950PRbatch_mat_mul_v332147471B->M->N线性映射+ASW滑窗蛇形(W=4)0117617610242561双缓冲驻留当前tile输入17617680allocate(输入驻留L2吸收重复读)direct_gm(输出直写GM不占L2)40方案B5252152522139True2L2场景B_输入驻留输出直写GM, r_in=1.00; 尾轮: 周长型主导, rho=0.06<rho_dv=0.53, A1b被dValue卡死, 方案B反超67108864.00.04.194304e-050.04.194304e-050.00.0274877906944.00.00056559240111934162684354560.000167772160.00.00056559240111934160.00.0005655924011193416MMADTrue计算BoundASW_Basic兜底 (StreamK未过: 1_并行缺口: P=B*MN*4B/L0C <= C/2; 2_单核K段下限: K/grid_K >= 256B/dtype 且 grid_K>=2; 3_归约代价可接受)
11asw_demo_reduce_core1616256256128bf16bf16bf16FalseFalseFalseTrue0asw_demo_reduce_coreASW_Basic_降核Ascend950PRbatch_mat_mul_v3161111降核: 只用16核, 每核一个L0C满载输出块, 其余核闲置012562561281281标准核内流水25625664allocatedirect_gm00不涉及(每核一块无尾轮)1110000True2P=16.00<C, 降核是理性选择 (强切则 tile 跌破搬移效率下限反而更慢)2097152.00.02.62144e-060.02.62144e-060.00.0268435456.01.104672658436214e-0620971522.62144e-060.02.62144e-060.02.62144e-06MTE2_GMTrue访存Bound(GM)ASW_Basic兜底 (StreamK未过: 2_单核K段下限: K/grid_K >= 256B/dtype 且 grid_K>=2)