Zahorian Stephen A, Hu Hongbing
Department of Electrical and Computer Engineering, State University of New York at Binghamton, Binghamton, New York 13902, USA.
J Acoust Soc Am. 2008 Jun;123(6):4559-71. doi: 10.1121/1.2916590.
In this paper, a fundamental frequency (F(0)) tracking algorithm is presented that is extremely robust for both high quality and telephone speech, at signal to noise ratios ranging from clean speech to very noisy speech. The algorithm is named "YAAPT," for "yet another algorithm for pitch tracking." The algorithm is based on a combination of time domain processing, using the normalized cross correlation, and frequency domain processing. Major steps include processing of the original acoustic signal and a nonlinearly processed version of the signal, the use of a new method for computing a modified autocorrelation function that incorporates information from multiple spectral harmonic peaks, peak picking to select multiple F(0) candidates and associated figures of merit, and extensive use of dynamic programming to find the "best" track among the multiple F(0) candidates. The algorithm was evaluated by using three databases and compared to three other published F(0) tracking algorithms by using both high quality and telephone speech for various noise conditions. For clean speech, the error rates obtained are comparable to those obtained with the best results reported for any other algorithm; for noisy telephone speech, the error rates obtained are lower than those obtained with other methods.
本文提出了一种基频(F(0))跟踪算法,该算法在从清晰语音到非常嘈杂语音的信噪比范围内,对高质量语音和电话语音都具有极强的鲁棒性。该算法名为“YAAPT”,即“另一种基音跟踪算法”。该算法基于时域处理(使用归一化互相关)和频域处理的结合。主要步骤包括对原始声学信号和信号的非线性处理版本进行处理,使用一种新方法来计算包含多个频谱谐波峰值信息的修正自相关函数,进行峰值提取以选择多个F(0)候选值及相关品质因数,并广泛使用动态规划在多个F(0)候选值中找到“最佳”轨迹。该算法通过使用三个数据库进行评估,并在各种噪声条件下,将高质量语音和电话语音与其他三种已发表的F(0)跟踪算法进行比较。对于清晰语音,所获得的错误率与其他任何算法所报告的最佳结果相当;对于有噪声的电话语音,所获得的错误率低于其他方法。