딥페이크 검출을 위한 일반화된 메타러닝 EfficientNet 비전 변환기 모델
Generalized Meta-Learning EfficientNet Vision Transformer Model for Deepfake Detection
- 발행
- 2024 등록KCI API에는 발행일 항목이 없어 원문 등록일을 쓰고 있습니다. 실측으로 발행보다 최대 6년 늦습니다 — 재수집하면 발행연월로 바뀝니다.
- 소속·발행
- 국립부경대학교
- 출처
- 국내 KCI
- DOI
- 10.9717/kmms.2024.27.6.663
- 원문
- 원문 보기 ↗
개념
키워드
Deepfake Detection, Vision Transformer, Generalization, Video Forensics, Meta-Learning, EfficientNet, Deepfake Detection, Vision Transformer, Generalization, Video Forensics, Meta-Learning, EfficientNet
초록
Digitally manipulated images that are realistic-looking but fake, which are known as Deepfake. With the remarkable developments in deep generative models, the accessibility and accuracy of manipulated technologies are increasing, leading to fake videos becoming increasingly difficult to identify. Different facial forgery techniques result in complicated data distributions, but Deepfake detection techniques based on CNN(convolutional neural network) architecture are utilized in the majority of Deepfake de tection models as binary classification problems. In this paper, we propose a model, named MEViT, which uses a combination of EfficientNet Vision Transformer with a meta-learning-based technique to improve the generalization of the detection model. Furthermore, we propose a learning process to update the model and introduce pair-discrimination loss and domain adjustment loss to improve detection ability across various domains. We also create various experiments on several Deepfake datasets and com pare our proposal with many state-of-the-art works to prove the efficiency of our approach.