한국어 딥페이크 텍스트 탐지를 위한 BERT 계열 모델 성능 비교 및 도메인 전이 분석
Comparative Evaluation of BERT-Based Models for Korean Deepfake Text Detection with Domain Transfer Analysis
- 발행
- 2026 등록KCI API에는 발행일 항목이 없어 원문 등록일을 쓰고 있습니다. 실측으로 발행보다 최대 6년 늦습니다 — 재수집하면 발행연월로 바뀝니다.
- 소속·발행
- 제주대
- 출처
- 국내 KCI
- DOI
- 10.37272/JIECR.2026.2.26.1.243
- 원문
- 원문 보기 ↗
개념
키워드
Generative AI generated text, Deepfake text detection, Korean BERT-based models, Cross-domain generalization, Domain dependence, Generative AI generated text, Deepfake text detection, Korean BERT-based models, Cross-domain generalization, Domain dependence
초록
The rapid proliferation of generative artificial intelligence (AI) has led to the widespread production of AI-generated texts that are fluent and persuasive, yet potentially prone to factual distortion and reduced information reliability. To address these concerns, this study aims to develop a Korean-based deepfake text detector and to analyze its generalization performance across domains. Specifically, we collected and curated human- and AI-generated texts from two heterogeneous domains: tourist reviews of Seongsan Ilchulbong and YouTube comments related to youth employment. Based on these datasets, four training test combinations were designed, including two within-domain settings and two cross-domain settings. The detector was implemented by fine-tuning five Korean BERT-based models KoBERT, KoELECTRA, KcELECTRA, KLUE-BERT, and KLUE-RoBERTa under identical experimental conditions. Model performance was evaluated using accuracy, precision, recall, and F1-score. The experimental results indicate that all models achieved high performance in within-domain settings. However, cross-domain evaluation resulted in a 10 15% decrease in accuracy and F1-score, highlighting the strong doma