Abstractive Summarization of Long Educational Transcripts in the Era of GenAI
- 1 Department of Computer Science, Christ University, Bangalore, India
Abstract
The surge in online video data presents the imperative requirement of summarizing it for information compression, efficient understanding, and to support decision-making. Recent research studies have predominantly utilized structured input data for summarization. This paper explores the current state-of-the-art models in summarizing unstructured, conversational, and long transcript data. We further apply a transfer learning approach using a fine-tuned BART model, pre-trained on the SAMSum dataset. Hyperparameter tuning and selective layer freezing are applied to optimise model performance. This study focuses on abstractive summarization of long video transcript data. Our methodology integrates ChatGPT to generate reference abstractive summaries from long transcript data. A semi-automated pipeline using TextRank is proposed for reference summary generation. The proposed fine-tuned model shows a significant increase in ROUGE scores over the baseline model. The findings suggest that the transfer learning approach is effective for abstractive summarization of long transcript data in real-world conversational domains.
DOI: https://doi.org/10.3844/jcssp.2026.2652.2664
Copyright: © 2026 Deepa Fernandes and Rupali Sunil Wagh. This is an open access article distributed under the terms of the
Creative Commons Attribution License, which permits unrestricted use, distribution, and reproduction in any medium, provided the original author and source are credited.
- 31 Views
- 4 Downloads
- 0 Citations
Download
Keywords
- NLP
- Abstractive Text Summarization
- Transfer Learning
- BART
- Long Transcripts