ZETA (Zonal Encoding for Transformer Adaptation) : Leveraging Transformer Models, RAG, and Graph-Based Approaches for Multilingual Video Summarization
摘要
ZETA (Zonal Encoding for Transformer Application) is a multilingual Video Summarization Tool. This tool combines transformer models, retrieval-augmented generation (RAG), and graph-based keyword extraction to produce accurate summaries. ZETA uses the following components: Whisper for multilingual transcription, here we are able to generate both abstractive and extractive summarization. Here, the main advantage is the usage of RAG, which assists by improving speed and accuracy by extracting relevant transcript sections and key topics.