|
演講日期 :2026-10-27 •在大型語言模型(LLMs)的微調(Fine-tuning)架構中,聯邦學習(Federated Learning, FL)去中心化的特性使其極易受到資料中毒攻擊(Data Poisoning Attacks)的嚴峻威脅。惡意客戶端(Malicious Clients)可藉由在本地訓練階段注入有害行為資料集,隱蔽地操控模型。現有的防禦範式大多依賴於模型參數或梯度層面的異常偵測,然而,此類方法在高維度的參數空間中,難以有效捕捉並解構語意層面(Semantic-level)的隱蔽偏差。
•針對此技術瓶頸,本研究提出一套基於模型輸出行為分析(Behavior-based Output Analysis)的防禦框架。所提方法以有害分數(Harmful Score, HS)作為核心量化指標,透過評估各客戶端模型所生成回應的有害程度來精準識別攻擊。
Abstract •In the fine-tuning architecture of Large Language Models (LLMs), the decentralized nature of Federated Learning (FL) makes it highly vulnerable to the severe threat of data poisoning attacks. Malicious clients can covertly manipulate the model by injecting harmful behavioral datasets during the local training phase. Most existing defense paradigms primarily rely on anomaly detection at the model-parameter or gradient level; however, such approaches struggle to effectively capture and deconstruct hidden semantic-level deviations within high-dimensional parameter spaces.
•To address this technical bottleneck, this study proposes a defense framework based on model output behavior analysis. The proposed method adopts the Harmful Score (HS) as its core quantitative indicator, precisely identifying attacks by evaluating the degree of harmfulness in the responses generated by each client model.
|