隐私攻击的工程实现与防御

admin 2026-09-08 04:16:39 网络安全文章 来源:ZONE.CI 全球网 0 阅读模式

文章总结: 本文介绍了隐私攻击的工程实现与防御,重点涵盖成员推断攻击的三种方法(损失阈值法、熵法、参考模型法)及评估指标,训练数据提取攻击的前缀构造与置信度过滤,以及差分隐私的工程落地。建议采用差分隐私等防御措施降低隐私泄露风险。 综合评分: 85 文章分类: ai安全,数据安全


隐私攻击的工程实现与防御

原创

pandazhengzheng pandazhengzheng

安全分析与研究

2026年9月7日 22:00 广东

在小说阅读器读本章

去阅读

在公众号小说中沉浸阅读

一、成员推断攻击(MIA)工程实现

MIA判断某个样本是否在模型训练集中,是隐私泄露的基础度量。工程实现有三种主流方法。

1.1 损失阈值法

最简单的MIA:训练集样本的损失通常低于非训练集样本(过拟合)。

import numpy as np

class LossThresholdMIA:
    def __init__(self, model, threshold=None):
        self.model = model
        self.threshold = threshold

    def calibrate(self, known_members, known_nonmembers):
        """用已知成员/非成员标定阈值"""
        member_losses = [self._loss(x, y) for x, y in known_members]
        nonmember_losses = [self._loss(x, y) for x, y in known_nonmembers]
        # 选使分类准确率最高的阈值
        all_losses = member_losses + nonmember_losses
        best_t, best_acc = 0, 0
        for t in sorted(all_losses):
&nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; acc = np.mean([l < t&nbsp;for&nbsp;l&nbsp;in&nbsp;member_losses]) + \
&nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; np.mean([l >= t&nbsp;for&nbsp;l&nbsp;in&nbsp;nonmember_losses])
&nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp;&nbsp;if&nbsp;acc > best_acc:
&nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; best_t, best_acc = t, acc
&nbsp; &nbsp; &nbsp; &nbsp; self.threshold = best_t

&nbsp; &nbsp;&nbsp;def&nbsp;attack(self, x, y):
&nbsp; &nbsp; &nbsp; &nbsp;&nbsp;return&nbsp;self._loss(x, y) < self.threshold

&nbsp; &nbsp;&nbsp;def&nbsp;_loss(self, x, y):
&nbsp; &nbsp; &nbsp; &nbsp;&nbsp;return&nbsp;-np.log(self.model.predict_proba(x)[y] +&nbsp;1e-8)

1.2 熵法

熵法比损失法更鲁棒:训练集样本的预测熵通常更低(模型更”确信”)。

class&nbsp;EntropyMIA:
&nbsp; &nbsp;&nbsp;def&nbsp;__init__(self, model, threshold=None, augment_times=10):
&nbsp; &nbsp; &nbsp; &nbsp; self.model = model
&nbsp; &nbsp; &nbsp; &nbsp; self.threshold = threshold
&nbsp; &nbsp; &nbsp; &nbsp; self.augment_times = augment_times

&nbsp; &nbsp;&nbsp;def&nbsp;compute_modified_entropy(self, x, y):
&nbsp; &nbsp; &nbsp; &nbsp;&nbsp;"""对输入做数据增强后取熵的均值(更鲁棒)"""
&nbsp; &nbsp; &nbsp; &nbsp; entropies = []
&nbsp; &nbsp; &nbsp; &nbsp;&nbsp;for&nbsp;_&nbsp;in&nbsp;range(self.augment_times):
&nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; x_aug = self._augment(x)
&nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; probs = self.model.predict_proba(x_aug)
&nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; entropy = -np.sum(probs * np.log(probs +&nbsp;1e-8))
&nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; entropies.append(entropy)
&nbsp; &nbsp; &nbsp; &nbsp;&nbsp;return&nbsp;np.mean(entropies)

&nbsp; &nbsp;&nbsp;def&nbsp;attack(self, x, y):
&nbsp; &nbsp; &nbsp; &nbsp;&nbsp;return&nbsp;self.compute_modified_entropy(x, y) < self.threshold

1.3 参考模型法(Reference-based MIA)

用与目标模型同架构、在独立数据上训练的参考模型做相对损失比较:

class&nbsp;ReferenceMIA:
&nbsp; &nbsp;&nbsp;def&nbsp;__init__(self, target_model, reference_models):
&nbsp; &nbsp; &nbsp; &nbsp; self.target = target_model
&nbsp; &nbsp; &nbsp; &nbsp; self.references = reference_models

&nbsp; &nbsp;&nbsp;def&nbsp;attack(self, x, y):
&nbsp; &nbsp; &nbsp; &nbsp; target_loss = self._loss(self.target, x, y)
&nbsp; &nbsp; &nbsp; &nbsp; ref_losses = [self._loss(ref, x, y)&nbsp;for&nbsp;ref&nbsp;in&nbsp;self.references]
&nbsp; &nbsp; &nbsp; &nbsp; ref_mean = np.mean(ref_losses)
&nbsp; &nbsp; &nbsp; &nbsp;&nbsp;# 训练集样本:target_loss显著低于reference
&nbsp; &nbsp; &nbsp; &nbsp;&nbsp;return&nbsp;target_loss < ref_mean - self.margin

&nbsp; &nbsp;&nbsp;def&nbsp;_loss(self, model, x, y):
&nbsp; &nbsp; &nbsp; &nbsp;&nbsp;return&nbsp;-np.log(model.predict_proba(x)[y] +&nbsp;1e-8)

1.4 攻击效果评估

class&nbsp;MIAEvaluator:
&nbsp; &nbsp;&nbsp;def&nbsp;evaluate(self, attacker, members, nonmembers):
&nbsp; &nbsp; &nbsp; &nbsp; tp = sum(attacker.attack(x, y)&nbsp;for&nbsp;x, y&nbsp;in&nbsp;members)
&nbsp; &nbsp; &nbsp; &nbsp; fp = sum(attacker.attack(x, y)&nbsp;for&nbsp;x, y&nbsp;in&nbsp;nonmembers)
&nbsp; &nbsp; &nbsp; &nbsp; tn = len(nonmembers) - fp
&nbsp; &nbsp; &nbsp; &nbsp; fn = len(members) - tp
&nbsp; &nbsp; &nbsp; &nbsp; advantage = (tp / len(members)) + (tn / len(nonmembers)) -&nbsp;1
&nbsp; &nbsp; &nbsp; &nbsp;&nbsp;return&nbsp;{
&nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp;&nbsp;"accuracy": (tp + tn) / (len(members) + len(nonmembers)),
&nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp;&nbsp;"advantage": advantage, &nbsp;# MIA优势,0为无泄露
&nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp;&nbsp;"tpr": tp / len(members),
&nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp;&nbsp;"fpr": fp / len(nonmembers),
&nbsp; &nbsp; &nbsp; &nbsp; }

二、训练数据提取攻击

对生成式LLM,攻击者构造前缀让模型生成训练数据中的具体内容(如PII)。

2.1 前缀构造策略

class&nbsp;DataExtractionAttacker:
&nbsp; &nbsp;&nbsp;def&nbsp;__init__(self, target_llm, max_tokens=100):
&nbsp; &nbsp; &nbsp; &nbsp; self.llm = target_llm
&nbsp; &nbsp; &nbsp; &nbsp; self.max_tokens = max_tokens

&nbsp; &nbsp;&nbsp;def&nbsp;extract_with_prefix(self, prefix, n_samples=10):
&nbsp; &nbsp; &nbsp; &nbsp;&nbsp;"""用给定前缀生成候选提取文本"""
&nbsp; &nbsp; &nbsp; &nbsp; candidates = []
&nbsp; &nbsp; &nbsp; &nbsp;&nbsp;for&nbsp;_&nbsp;in&nbsp;range(n_samples):
&nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; output = self.llm.generate(
&nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; prompt=prefix, max_tokens=self.max_tokens, temperature=0.7
&nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; )
&nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; candidates.append(output)
&nbsp; &nbsp; &nbsp; &nbsp;&nbsp;return&nbsp;candidates

&nbsp; &nbsp;&nbsp;def&nbsp;extract_with_patterns(self, patterns):
&nbsp; &nbsp; &nbsp; &nbsp;&nbsp;"""用常见文档模式作为前缀(如"姓名:""邮箱:")"""
&nbsp; &nbsp; &nbsp; &nbsp; results = {}
&nbsp; &nbsp; &nbsp; &nbsp;&nbsp;for&nbsp;pattern&nbsp;in&nbsp;patterns:
&nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; candidates = self.extract_with_prefix(pattern)
&nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; results[pattern] = self._filter_sensitive(candidates)
&nbsp; &nbsp; &nbsp; &nbsp;&nbsp;return&nbsp;results

2.2 置信度过滤

模型对训练数据的生成通常置信度更高,用置信度过滤候选:

class&nbsp;ConfidenceFilter:
&nbsp; &nbsp;&nbsp;def&nbsp;__init__(self, llm, threshold=0.8):
&nbsp; &nbsp; &nbsp; &nbsp; self.llm = llm
&nbsp; &nbsp; &nbsp; &nbsp; self.threshold = threshold

&nbsp; &nbsp;&nbsp;def&nbsp;filter(self, candidates):
&nbsp; &nbsp; &nbsp; &nbsp; filtered = []
&nbsp; &nbsp; &nbsp; &nbsp;&nbsp;for&nbsp;text&nbsp;in&nbsp;candidates:
&nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; logprobs = self.llm.get_logprobs(text)
&nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; avg_logprob = np.mean(logprobs)
&nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp;&nbsp;# 转为概率
&nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; avg_prob = np.exp(avg_logprob)
&nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp;&nbsp;if&nbsp;avg_prob > self.threshold:
&nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; filtered.append({"text": text,&nbsp;"confidence": avg_prob})
&nbsp; &nbsp; &nbsp; &nbsp;&nbsp;return&nbsp;filtered

2.3 提取验证

判断提取出的内容是否真的是训练数据(而非模型”幻觉”):

class&nbsp;ExtractionVerifier:
&nbsp; &nbsp;&nbsp;def&nbsp;__init__(self, llm, canary_db=None):
&nbsp; &nbsp; &nbsp; &nbsp; self.llm = llm
&nbsp; &nbsp; &nbsp; &nbsp; self.canary_db = canary_db &nbsp;# 已知训练数据(如canary插入)

&nbsp; &nbsp;&nbsp;def&nbsp;verify(self, extracted_text):
&nbsp; &nbsp; &nbsp; &nbsp;&nbsp;# 1. 与已知canary比对
&nbsp; &nbsp; &nbsp; &nbsp;&nbsp;if&nbsp;self.canary_db&nbsp;and&nbsp;extracted_text&nbsp;in&nbsp;self.canary_db:
&nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp;&nbsp;return&nbsp;{"verified":&nbsp;True,&nbsp;"method":&nbsp;"canary_match"}
&nbsp; &nbsp; &nbsp; &nbsp;&nbsp;# 2. 一致性检查:多次生成是否稳定复现
&nbsp; &nbsp; &nbsp; &nbsp; reproductions = [
&nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; self.llm.generate(prompt=extracted_text[:20], temperature=0)
&nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp;&nbsp;for&nbsp;_&nbsp;in&nbsp;range(5)
&nbsp; &nbsp; &nbsp; &nbsp; ]
&nbsp; &nbsp; &nbsp; &nbsp; consistency = np.mean([
&nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; r.startswith(extracted_text[:50])&nbsp;for&nbsp;r&nbsp;in&nbsp;reproductions
&nbsp; &nbsp; &nbsp; &nbsp; ])
&nbsp; &nbsp; &nbsp; &nbsp;&nbsp;return&nbsp;{"verified": consistency >&nbsp;0.8,&nbsp;"method":&nbsp;"consistency"}

三、差分隐私工程落地

DP提供形式化隐私保证:相邻数据集(差一个样本)的输出分布不可区分。

3.1 DP-SGD代码实现

import&nbsp;torch

`


免责声明:

本文所载程序、技术方法仅面向合法合规的安全研究与教学场景,旨在提升网络安全防护能力,具有明确的技术研究属性。

任何单位或个人未经授权,将本文内容用于攻击、破坏等非法用途的,由此引发的全部法律责任、民事赔偿及连带责任,均由行为人独立承担,本站不承担任何连带责任。

本站内容均为技术交流与知识分享目的发布,若存在版权侵权或其他异议,请通过邮件联系处理,具体联系方式可点击页面上方的联系我。

本文转载自:安全分析与研究 pandazhengzheng pandazhengzheng《隐私攻击的工程实现与防御》

隐私攻击的工程实现与防御 网络安全文章

隐私攻击的工程实现与防御

文章总结: 本文介绍了隐私攻击的工程实现与防御,重点涵盖成员推断攻击的三种方法(损失阈值法、熵法、参考模型法)及评估指标,训练数据提取攻击的前缀构造与置信度过滤
评论:0   参与:  0