Files
matmul-analysis/BMM/BMM_Theory/examples/result_evaluate.csv

6.9 KiB

1case_idbatch_abatch_bmnkdtype_adtype_bdtype_ctrans_atrans_bhas_biasout_nddeterministic_levelplan_case_idplan_branchplan_npuplan_opplan_used_core_numplan_split_bplan_m_cntplan_n_cntplan_grid_kplan_core_mapplan_b_coreplan_merge_b0plan_single_core_mplan_single_core_nplan_single_core_kplan_k_l1plan_b_l1plan_l1_formplan_base_mplan_base_nplan_base_kplan_l2_policy_inplan_l2_policy_outplan_swizzle_wplan_workspace_bytesplan_tail_strategyplan_fixpipe_unitflagplan_out_dtype_bytesplan_notegm_read_bytesl2_read_bytest_mte2_gmt_mte2_l2t_mte2dma_cmd_countt_dma_cmdcube_flopst_mmadfixpipe_bytest_fixpipet_reducet_steadyt_draint_totalbottleneckfeasibleviolationsbound_typeadvice
2merge_demo_11281286464512bf16bf16bf16FalseFalseFalseTrue0merge_demo_1IterBatchAscend950PRbatch_mat_mul_v33232111切B轮转分配(核间零重复读零依赖)4164645125122b_双batch乒乓6464256allocate(GM->L1随路驻留L2)direct_gm(输出仅写一次,直写GM不占L2)00不涉及(核内不切M/N)True2双batch乒乓: 2*(MK+KN)*dtype=256KB <= L15242880.01.048576e-050.01.0685759999999999e-0542.0000000000000002e-0716777216.01.104672658436214e-06327686.5536e-070.01.0685759999999999e-054.400081646090535e-071.1125768164609053e-05MTE2_GMTrue访存Bound(GM)瓶颈在 GM 搬入: 可考虑增大 tile 提升 dValue/单核搬移量, 或利用 L2 驻留吸收重复读 (MergeBatch/ASW swizzle 方向)
3iter_demo_11281286464256bf16bf16bf16FalseFalseFalseTrue0iter_demo_1IterBatchAscend950PRbatch_mat_mul_v33232111切B轮转分配(核间零重复读零依赖)4164642562562b_双batch乒乓6464256allocate(GM->L1随路驻留L2)direct_gm(输出仅写一次,直写GM不占L2)00不涉及(核内不切M/N)True2双batch乒乓: 2*(MK+KN)*dtype=128KB <= L12621440.05.24288e-060.05.4428799999999995e-0642.0000000000000002e-078388608.05.52336329218107e-07327686.5536e-070.05.4428799999999995e-063.0192408230452674e-075.744804082304526e-06MTE2_GMTrue访存Bound(GM)瓶颈在 GM 搬入: 可考虑增大 tile 提升 dValue/单核搬移量, 或利用 L2 驻留吸收重复读 (MergeBatch/ASW swizzle 方向)
4iter_demo_2512512128128128bf16bf16bf16FalseFalseFalseTrue0iter_demo_2IterBatchAscend950PRbatch_mat_mul_v33232111切B轮转分配(核间零重复读零依赖)1611281281281282b_双batch乒乓128128128allocate(GM->L1随路驻留L2)direct_gm(输出仅写一次,直写GM不占L2)00不涉及(核内不切M/N)True2双batch乒乓: 2*(MK+KN)*dtype=128KB <= L110485760.02.097152e-050.02.1771519999999998e-05168.000000000000001e-0767108864.04.418690633744856e-065242881.048576e-050.02.1771519999999998e-059.315281646090535e-072.270304816460905e-05MTE2_GMTrue访存Bound(GM)瓶颈在 GM 搬入: 可考虑增大 tile 提升 dValue/单核搬移量, 或利用 L2 驻留吸收重复读 (MergeBatch/ASW swizzle 方向)
5iter_demo_32048204810241024512bf16bf16bf16FalseFalseFalseTrue0iter_demo_3IterBatchAscend950PRbatch_mat_mul_v33232111切B轮转分配(核间零重复读零依赖)64110241024512641d_两侧都切K32102416allocate(GM->L1随路驻留L2)direct_gm(输出仅写一次,直写GM不占L2)00不涉及(核内不切M/N)True2两侧都切K: k_L1=64, K段成对流水, batch边界天然无缝1342177280.00.002684354560.00.002709954565122.5600000000000002e-0568719476736.00.0045247392089547331342177280.002684354560.00.0045247392089547335.078042126748971e-050.004575519630222223MMADTrue计算Bound瓶颈在 Cube 计算: 已接近理论算力上限, 检查是否有冗余计算 (MergeBatch 交叉项) 可消除
6iter_demo_4646464648192bf16bf16bf16FalseFalseFalseTrue0iter_demo_4IterBatchAscend950PRbatch_mat_mul_v33232111切B轮转分配(核间零重复读零依赖)216464819210241d_两侧都切K6464256allocate(GM->L1随路驻留L2)direct_gm(输出仅写一次,直写GM不占L2)00不涉及(核内不切M/N)True2两侧都切K: k_L1=1024, K段成对流水, batch边界天然无缝41943040.08.388608e-050.08.468608e-05168.000000000000001e-07134217728.08.837381267489712e-06163843.2768e-070.08.468608e-057.16176329218107e-078.540225632921811e-05MTE2_GMTrue访存Bound(GM)瓶颈在 GM 搬入: 可考虑增大 tile 提升 dValue/单核搬移量, 或利用 L2 驻留吸收重复读 (MergeBatch/ASW swizzle 方向)
7fp16_out_demo1281286464512bf16bf16fp16FalseFalseFalseTrue0fp16_out_demoIterBatchAscend950PRbatch_mat_mul_v33232111切B轮转分配(核间零重复读零依赖)4164645125122b_双batch乒乓6464256allocate(GM->L1随路驻留L2)direct_gm(输出仅写一次,直写GM不占L2)00不涉及(核内不切M/N)True2双batch乒乓: 2*(MK+KN)*dtype=256KB <= L15242880.01.048576e-050.01.0685759999999999e-0542.0000000000000002e-0716777216.01.104672658436214e-06327686.5536e-070.01.0685759999999999e-054.400081646090535e-071.1125768164609053e-05MTE2_GMTrue访存Bound(GM)瓶颈在 GM 搬入: 可考虑增大 tile 提升 dValue/单核搬移量, 或利用 L2 驻留吸收重复读 (MergeBatch/ASW swizzle 方向)
8to_matmul_demo11204820482048bf16bf16bf16FalseFalseFalseTrue0to_matmul_demo转MatmulAscend950PRbatch_mat_mul_v3321111010000100000True2BatchA=1或BatchB=1, 折叠转普通Matmul (该分支详实现待后续迭代)0.00.00.00.00.00.00.00.00.00.00.00.00.00.00.0FalsedValue=0B < 下限 128B, 搬移带宽利用率崩塌分支 '转Matmul' 的评估模型待后续迭代; 当前支持: ['IterBatch', 'MergeBatch']
9special_k11281282562561bf16bf16bf16FalseFalseFalseTrue0special_k1特殊分支Ascend950PRbatch_mat_mul_v3321111010000100000True2K=1, Cube 无用, 走 AIV 向量通路 (该分支详实现待后续迭代)0.00.00.00.00.00.00.00.00.00.00.00.00.00.00.0FalsedValue=0B < 下限 128B, 搬移带宽利用率崩塌分支 '特殊分支' 的评估模型待后续迭代; 当前支持: ['IterBatch', 'MergeBatch']
10asw_fallback_demo3232409640964096bf16bf16bf16FalseFalseFalseTrue0asw_fallback_demoASW_BasicAscend950PRbatch_mat_mul_v3321111010000100000True2IterBatch/MergeBatch 进入条件均不满足 (如负载均衡/搬移效率不达标), 回落 ASW_Basic (该分支详实现待后续迭代); IterBatch未过: 3_L1驻留形态(四选一, 核心要求: 单batch核内零重复读); MergeBatch未过: 1_batch关系与每核份额: BatchA==BatchB 且 b_core=B/C>=2*b0; 2_L0C容量: 2*(b0*M)*(b0*N)*4B <= L0C; 5_访存Bound: 2MN/(M+N) < R16/b00.00.00.00.00.00.00.00.00.00.00.00.00.00.00.0FalsedValue=0B < 下限 128B, 搬移带宽利用率崩塌分支 'ASW_Basic' 的评估模型待后续迭代; 当前支持: ['IterBatch', 'MergeBatch']