Files
matmul-analysis/BMM/BMM_Theory/t_out.csv
2026-09-03 12:37:16 +00:00

3.0 KiB

1case_idbatch_abatch_bmnkdtype_adtype_bdtype_ctrans_atrans_bhas_biasout_nddeterministic_levelplan_case_idplan_branchplan_npuplan_opplan_used_core_numplan_split_bplan_m_cntplan_n_cntplan_grid_kplan_core_mapplan_b_coreplan_merge_b0plan_single_core_mplan_single_core_nplan_single_core_kplan_k_l1plan_b_l1plan_l1_formplan_base_mplan_base_nplan_base_kplan_l2_policy_inplan_l2_policy_outplan_swizzle_wplan_workspace_bytesplan_tail_strategyplan_tail_m_cntplan_tail_n_cntplan_tail_k_cntplan_tail_m_mainplan_tail_n_mainplan_tail_block_cntplan_tail_wave_numplan_fixpipe_unitflagplan_out_dtype_bytesplan_notegm_read_bytesl2_read_bytest_mte2_gmt_mte2_l2t_mte2dma_cmd_countt_dma_cmdcube_flopst_mmadfixpipe_bytest_fixpipet_reducet_steadyt_draint_totalbottleneckfeasibleviolationsbound_typeadvice
2x515125542682fp32fp32fp32FalseFalseFalseTrue0xStreamKAscend950PRbatch_mat_mul_v3321113B/M/N切出51块, 每块3核切K归约 (归约组内核c负责K段[c*K/3,(c+1)*K/3))11255422281281K段标准分块流水1284264allocate(部分和驻留L2)resident(部分和4B驻留L2, 防精度丢失不随C的fp16/fp8转换)06554520grid_K=3路切K+归约1130000True4P=8.33, grid_K=3, 部分和驻留L2按4B写出, AIV归约后按C dtype=4B写最终465578.66666666670.09.311573333333333e-060.09.311573333333333e-060.00.029797034.6666666681.9619446694101507e-062621441.6131938461538462e-063.674316083916084e-079.311573333333333e-063.674316083916084e-079.679004941724942e-06MTE2_GMTrue访存Bound(GM)切B分支条件不满足, 落 StreamK
3y22641632fp16fp16fp16FalseFalseFalseTrue0yASW_Basic_降核Ascend950PRbatch_mat_mul_v311111降核: 只用1核, 每核一个L0C满载输出块, 其余核闲置01641632321标准核内流水641632allocatedirect_gm00不涉及(每核一块无尾轮)1110000True2P=0.03<C, 降核是理性选择 (强切则 tile 跌破搬移效率下限反而更慢)10240.00.02.048e-070.02.048e-070.00.0131072.08.630255144032922e-0940968.192e-080.02.048e-070.02.048e-07MTE2_GMTrue访存Bound(GM)ASW_Basic兜底 (StreamK未过: 2_单核K段下限: K/grid_K >= 256B/dtype 且 grid_K>=2; 3_归约代价可接受: K > grid_K^2/(grid_K-1)*theta_c) [自检违规: dValue=64B < 下限 128B, K 段连续维搬移效率崩塌] —— 方案生成存在缺陷, 需人工复核
4z25625612810248bf16bf16bf16FalseFalseFalseTrue0zIterBatchAscend950PRbatch_mat_mul_v33232111切B轮转分配(核间零重复读零依赖)811281024882b_双batch乒乓12825616allocate(GM->L1随路驻留L2)direct_gm(输出仅写一次,直写GM不占L2)00不涉及(核内不切M/N)1110000True2双batch乒乓: 2*(MK+KN)*dtype=36KB <= L11474560.02.94912e-060.03.34912e-0684.0000000000000003e-0716777216.01.104672658436214e-0620971524.194304e-050.04.194304e-055.380964082304527e-064.7324004082304525e-05FIXPIPETrue写出Bound仅 IterBatch 条件满足