|
Description
|
As YouTube content continues to grow, advanced filtering systems are crucial to ensuring a safe and enjoyable user experience. We present MFusTSVD, a multi-modal model for classifying YouTube video content by analyzing text, audio, and video images. MFusTSVD uses specialized methods to extract features from audio and video images, while processing text data with BERT Transformers. Our key innovation includes two new BERT-based multi-modal fusion methods: B-SMTLMF and B-CMTLRMF. These methods combine features from different data types and improve the model's ability to understand each type of data, including detailed audio patterns, leading to better content classification and speech-related separation. MFusTSVD is designed to perform better than existing models in terms of accuracy, precision, recall, and F-measure. Tests show that MFusTSVD consistently outperforms popular models like Memory Fusion Network, Early Fusion LSTM, Late Fusion LSTM, and multi-modal Transformer across different content types and evaluation measures. In particular, MFusTSVD effectively balances precision and recall, which makes it especially useful for identifying inappropriate speech and audio content, as well as broader categories, ensuring reliable and robust content moderation. (2025-05-01)
***This entry has been automatically imported via OpenAlex by LIST harvest scripts. Please refer to https://doi.org/10.1109/jstsp.2025.3569446 for the original and latest version of the publication*** (2026-07-01)
|
|
Keyword
|
Computer science, Artificial intelligence, Computer vision, Transformer, Fusion, Sensor fusion, Content (measure theory), Pattern recognition (psychology), Speech recognition |