搜索

x
中国物理学会期刊

一种基于MPI的并行三维WCS-FDTD方法

An MPI-Based Parallel Three-Dimensional WCS-FDTD Method

PDF
导出引用
  • 针对两个方向存在精细结构的大规模三维电磁问题,弱条件稳定时域有限差分方法虽能放宽传统时域有限差分方法的时间步长限制,但需反复求解大规模三对角矩阵方程,单机计算能力和存储空间难以满足求解需求.为此,提出一种基于消息传递接口的并行弱条件稳定时域有限差分方法.该方法在精细网格方向采用隐式迭代,使时间步长主要由粗网格方向决定.同时,沿粗网格方向对计算区域进行一维划分,使各子区域内的三对角方程可独立求解,并通过相邻进程交换边界场量实现并行计算.以金属超光栅为算例,对近场探针结果,远场雷达散射截面和计算效率进行验证.结果表明,所提方法与传统FDTD方法及串行WCS-FDTD方法的计算结果吻合良好.在相同12进程条件下,并行WCS-FDTD的计算时间约为并行FDTD的1/4.2.在32进程下,加速比达到30.5,并行效率为95.3%.进一步的多节点测试表明,当每节点固定使用16个MPI进程,计算节点由1个增加至3个时,计算时间由223.2 s降低至89.4 s,相对于单节点16进程结果,加速比为2.5,跨节点并行效率为83.2%.该方法能够同时利用多个计算节点的计算与存储资源,具有较好的跨节点扩展能力,适用于两个方向含精细结构的大规模三维电磁问题.

     

    For large-scale three-dimensional electromagnetic problems with fine structures in two spatial directions, the Courant-Friedrichs-Lewy (CFL) condition imposes a severe restriction on the finite-difference time-domain (FDTD) method. Its time step is limited by the smallest grid size, resulting in low computational efficiency. The weakly conditionally stable FDTD (WCS-FDTD) method relaxes this restriction through implicit field updates. However, large tridiagonal systems must be solved repeatedly at each time step. These solutions impose substantial computational and memory demands on a single computing node. To overcome these limitations, an MPI-based parallel three-dimensional WCS-FDTD method is proposed for distributed-memory platforms. The field components along the two finely discretized directions are updated implicitly. Therefore, the time step is determined by the grid size in the coarsely discretized direction. The computational domain is decomposed one-dimensionally along this coarse-grid direction. As a result, each MPI process retains complete grid lines in both implicit directions. It independently solves the local tridiagonal systems using the Thomas algorithm. No interprocess communication is required during the implicit solutions. After the corresponding field updates, neighboring processes exchange only the required boundary field components through MPI Sendrecv. This strategy combines the enlarged time step of WCS-FDTD with process-level parallelism. It also distributes the computational workload and field data across multiple nodes. A metallic metagrating is used to evaluate the accuracy, computational efficiency, and parallel scalability of the proposed method. The calculated near-field probe responses agree closely with those obtained using the conventional FDTD and serial WCS-FDTD methods. Close agreement is also observed for the far-field radar cross sections. For this example, WCS-FDTD increases the time step from 0.70 ps to 6.67 ps. It reduces the serial runtime by a factor of 3.8 relative to conventional FDTD. When the same 12 MPI processes are used, parallel WCS-FDTD is 4.2 times faster than parallel FDTD. It is also 11.5 times faster than serial WCS-FDTD. With 32 processes, the parallel WCS-FDTD method achieves a speedup of 30.5 and a parallel efficiency of 95.3%. Multi-node tests are also performed using 16 MPI processes per node. Increasing the number of nodes from one to three reduces the runtime from 223.2 s to 89.4 s. This corresponds to a relative speedup of 2.5 and a cross-node parallel efficiency of 83.2%. These results demonstrate that the proposed method preserves the accuracy of WCS-FDTD and substantially improves computational efficiency. It also effectively exploits the aggregated computing and memory resources of multiple nodes. Therefore, the method is well suited to large-scale three-dimensional electromagnetic simulations involving fine structures in two spatial directions.

     

    目录

    /

    返回文章
    返回