Unlearn target PI
Modify the model on a designated forget set.
Private-information unlearning · White-box auditing
We study whether private information targeted by LLM unlearning is genuinely removed or merely suppressed under standard generation. White-box audits of output logits reveal that information which appears forgotten can remain recoverable.
Forgotten does not equal Erased.
The evaluation gap
Private information is often structured and uniquely defined. A model can stop producing a target value under ordinary decoding while still retaining a recoverable trace.
Modify the model on a designated forget set.
Standard decoding no longer returns the target PI.
Probe attribute-relevant logits for what remains.
Real private data should not be released for benchmarking. FPI instead builds fictitious profiles and paraphrased question–answer pairs, providing known ground truth for clean, attribute-specific auditing.
Each person–attribute pair has four paraphrased questions with the same answer. The experiments evaluate five finetuning-based methods: GA, GD, GA+KL, PO, and NPO.
View larger
The audit asks a stronger question than black-box evaluation: given the unlearned model and knowledge of how it was produced, how much target PI can an auditor recover?
White-box auditor
Access
Goal
The experiments focus on finetuning-based methods, where model parameters change. Prompt- and decoding-based approaches leave the underlying transformer unchanged and are straightforward to reverse in this white-box setting.
Reading forget quality. Higher forget quality means more error on the forget set under a given decoding strategy. If an audit makes that score fall, the model is again producing values closer to the ground truth—evidence of recovery.
Loss-increasing unlearning can move target tokens into low-likelihood regions rather than remove their trace. Restricted Inverse Greedy (RIG) tests this possibility by selecting the least likely token at each step within an attribute-valid candidate set.
Illustration · attribute-valid candidate space
Schematic relative likelihoods—not measured values
Restricted Greedy selects the highest-likelihood option inside the valid attribute space.
Restricted Inverse Greedy selects the lowest-likelihood option inside that same valid space.
RIG audits GA, GD, and GA+KL, which directly increase loss on the original forget data. Candidate sets encode valid forms: digits for SIN, alternating letters and digits for postcode, the known year range, or the eight blood-type labels.
RG is used for Preference Optimization behavior. PO lowers loss on modified question–rejection pairs rather than increasing loss on the original answer, so high-likelihood tokens within the valid attribute space are the appropriate probe. RIG and RG are not interchangeable.
Main result · DeepSeek-7B
Across attributes, standard decoding often indicates strong forgetting. RIG or RG can sharply lower the measured forget quality, showing that target information remains accessible through the model’s logits.
View full-size figure
Standard versus audited. The red bars show original forget quality; green and blue show RIG and RG audits. A lower audited score means predictions move closer to the true PI. PO is audited with RG and reported separately in the paper.
Why retraining matters. The retrained model f* never sees forget samples and does not show the same consistent audit-induced decline, providing the gold-standard reference for approximate unlearning.
Additional analyses
Supporting experiments reproduce the principal comparison on Qwen3-8B and trace how apparent forgetting and audited recovery diverge during GD unlearning.
View larger
The main original-versus-audited forget-quality comparison is reproduced on Qwen3-8B.
View larger
During GD training, standard forget quality rises while audited forget quality under RIG generally falls.
The paper finds a recurring tension: settings that resist RIG often incur severe utility degradation, while utility-preserving settings tend to leave recoverable traces. Each plot compares GA, GD, GA+KL, and NPO after 50 unlearning iterations.
For NPO on postcode and SIN, restricted beam search finds lower-error candidate sets than random guessing, showing that traces can remain even when the simpler RG probe finds little.
Project materials
Read the paper, inspect the official implementation, or use the FPI benchmark to reproduce and extend the study.
Citation
@misc{hu2026recoverability,
title={On the Recoverability of Private Information Unlearning in Large Language Models},
author={Shicheng Hu and Runzhi Tian and Ziqiao Wang and Yongyi Mao},
year={2026},
eprint={2608.29943},
archivePrefix={arXiv},
primaryClass={cs.LG},
url={https://arxiv.org/abs/2608.29943}
}