MFCC
2 articles filed under this keyword.
Kurdish Spoken Letter Recognition based on k-NN and SVM Model
Zrar Khalid Abdul
Automatic recognition of spoken letters is one of the most challenging tasks in the area of speech recognition system. In this paper, different machine learning approaches are used to classify the Kurdish alphabets such as SVM and k-NN where both approaches are fed by two different features, Linear Predictive Coding (LPC) and Mel Frequency Cepstral Coefficients (MFCCs). Moreover, the features are combined together to learn the classifiers. The experiments are evaluated on the dataset that are collected by the authors as there as not standard Kurdish dataset. The dataset consists of 2720 samples as a total. The results show that the MFCC features outperforms the LPC features as the MFCCs have more relative information of vocal track. Furthermore, fusion of the features (MFCC and LPC) is not capable to improve the classification rate significantly.
Volume 6 · Issue 2 · September 2019Uttered Kurdish digit recognition system
Saman Muhammad Omer, Jihad Anwar Qadir, Zrar Khalid Abdul
Speech recognition is a crucial subject in human computer interaction area. The ability of a machine to recognize words and phrases in spoken language is speech recognition and then convert them to a machine-readable format. Digit recognition is a part of the speech recognition system. In this paper, three spectral based features including Mel Frequency Cepstral Coefficient (MFCC), Linear predictive coding (LPC) and formant frequencies are proposed to classify ten Kurdish uttered digits (0-9). The features are extracted from entire speech signal, and feed a pairwise SVM classifier. Experiments including each individual feature and different forms of fusion are conducted and the results are shown. The fusion of the features significantly improves the result and shows that the different features carry complementary information. The proposed model is experimented on the dataset that have been collected in Kurdistan. Key words: Speech recognition, MFCC, LPC, Formant frequencies, uttered digits, SVM