EfficientNet와 Visual Transformer에 기반한 투 스트림 딥페이크 영상 탐지 알고리즘 제안
Proposing A Two-Stream Deepfake Video Detection Algorithm Based on EfficientNet and Visual Transformer
- 발행
- 2025 등록KCI API에는 발행일 항목이 없어 원문 등록일을 쓰고 있습니다. 실측으로 발행보다 최대 6년 늦습니다 — 재수집하면 발행연월로 바뀝니다.
- 소속·발행
- 한국외국어대학교
- 출처
- 국내 KCI
- DOI
- 10.9717/kmms.2025.28.3.470
- 원문
- 원문 보기 ↗
개념
키워드
Deep Learning, Computer Vision, Deep Fake, EfficientNetwork, Visual Transformer, Deep Learning, Computer Vision, Deep Fake, EfficientNetwork, Visual Transformer
초록
This paper proposes a model using a two-stream network to extract both global and local spatial forgery features in deepfake videos and learn temporal forgery features across frames. The spatial detection model is based on EfficientNet, and the temporal detection model is based on ViT. Spatial features are divided into two streams to extract global and local features separately, while two ViTs are used to effectively learn temporal features, making the temporal part also two-stream. This achieved an average 9.4% improvement in detection accuracy compared to existing deepfake detection models, and through Grad-CAM, we visually confirmed the regions that significantly influenced the determination of deepfakes, demonstrating the model's effective detection capabilities.