The research team proposed the SoundMHPE method for the first time, aiming to estimate multi-person 3D poses using only acoustic signals. The framework consists of two core components: the Acoustic Multi-scale Encoder, which separates subtle features from complex overlapping acoustic signals; and the Temporal Pose Decoder, which decouples cross-frame information through an attention mechanism to reconstruct individual poses. To verify the effectiveness of this method, the researchers constructed an AMP dataset containing 432,000 frames of synchronized data. Experimental results show that SoundMHPE outperforms existing baseline models in performance.