Files
matmul-analysis/BMM/BMM_Theory/t_p.csv

1.4 KiB

1case_idbranchnpuopused_core_numsplit_bm_cntn_cntgrid_kcore_mapb_coremerge_b0single_core_msingle_core_nsingle_core_kk_l1b_l1l1_formbase_mbase_nbase_kl2_policy_inl2_policy_outswizzle_wworkspace_bytestail_strategytail_m_cnttail_n_cnttail_k_cnttail_m_maintail_n_maintail_block_cnttail_wave_numfixpipe_unitflagout_dtype_bytesnote
2xStreamKAscend950PRbatch_mat_mul_v3321113B/M/N切出51块, 每块3核切K归约 (归约组内核c负责K段[c*K/3,(c+1)*K/3))11255422281281K段标准分块流水1284264allocate(部分和驻留L2)resident(部分和4B驻留L2, 防精度丢失不随C的fp16/fp8转换)06554520grid_K=3路切K+归约1130000True4P=8.33, grid_K=3, 部分和驻留L2按4B写出, AIV归约后按C dtype=4B写最终
3yASW_Basic_降核Ascend950PRbatch_mat_mul_v311111降核: 只用1核, 每核一个L0C满载输出块, 其余核闲置01641632321标准核内流水641632allocatedirect_gm00不涉及(每核一块无尾轮)1110000True2P=0.03<C, 降核是理性选择 (强切则 tile 跌破搬移效率下限反而更慢)
4zIterBatchAscend950PRbatch_mat_mul_v33232111切B轮转分配(核间零重复读零依赖)811281024882b_双batch乒乓12825616allocate(GM->L1随路驻留L2)direct_gm(输出仅写一次,直写GM不占L2)00不涉及(核内不切M/N)1110000True2双batch乒乓: 2*(MK+KN)*dtype=36KB <= L1