Private-information unlearning · White-box auditing

On the Recoverability of Private Information Unlearning in Large Language Models

Shicheng Hu† Runzhi Tian† Ziqiao Wang Yongyi Mao

† Equal contribution

We study whether private information targeted by LLM unlearning is genuinely removed or merely suppressed under standard generation. White-box audits of output logits reveal that information which appears forgotten can remain recoverable.

Forgotten does not equal Erased.

The evaluation gap

Forgotten—or merely hidden?

Private information is often structured and uniquely defined. A model can stop producing a target value under ordinary decoding while still retaining a recoverable trace.

1

Unlearn target PI

Modify the model on a designated forget set.

2

Appears forgotten

Standard decoding no longer returns the target PI.

3

Audit recoverability

Probe attribute-relevant logits for what remains.

Controlled evaluation

The Fake Private Information dataset

Real private data should not be released for benchmarking. FPI instead builds fictitious profiles and paraphrased question–answer pairs, providing known ground truth for clean, attribute-specific auditing.

400
Synthetic profiles
6,400
Question–answer pairs
4
Private attributes
16
QA pairs per profile
Year of birth Blood type Postcode Social insurance number

Each person–attribute pair has four paraphrased questions with the same answer. The experiments evaluate five finetuning-based methods: GA, GD, GA+KL, PO, and NPO.

Diagram showing fictitious personal profiles with year of birth, SIN, postcode, and blood type transformed into question–answer pairs. View larger
FPI construction: fictitious profiles are transformed into paraphrased question–answer pairs. Values shown in the figure are synthetic.

Threat model

White-box auditing

The audit asks a stronger question than black-box evaluation: given the unlearned model and knowledge of how it was produced, how much target PI can an auditor recover?

White-box auditor

Access

  • Unlearned model and output logits
  • Unlearning algorithm knowledge
  • Retain set
  • Forget-set prompts and, for FPI, masked answer text

Goal

  • Recover as much target PI as possible
  • Distinguish information erasure from output suppression

The experiments focus on finetuning-based methods, where model parameters change. Prompt- and decoding-based approaches leave the underlying transformer unchanged and are straightforward to reverse in this white-box setting.

Reading forget quality. Higher forget quality means more error on the forget set under a given decoding strategy. If an audit makes that score fall, the model is again producing values closer to the ground truth—evidence of recovery.

Recovery methods

Look where unlearning pushes the answer

Loss-increasing unlearning can move target tokens into low-likelihood regions rather than remove their trace. Restricted Inverse Greedy (RIG) tests this possibility by selecting the least likely token at each step within an attribute-valid candidate set.

Illustration · attribute-valid candidate space

Schematic relative likelihoods—not measured values

1975
1976
1987 ↑ RG
1998 ↑ RIG
2005
RG

Restricted Greedy selects the highest-likelihood option inside the valid attribute space.

RIG

Restricted Inverse Greedy selects the lowest-likelihood option inside that same valid space.

RIG audits GA, GD, and GA+KL, which directly increase loss on the original forget data. Candidate sets encode valid forms: digits for SIN, alternating letters and digits for postcode, the known year range, or the eight blood-type labels.

RG is used for Preference Optimization behavior. PO lowers loss on modified question–rejection pairs rather than increasing loss on the original answer, so high-likelihood tokens within the valid attribute space are the appropriate probe. RIG and RG are not interchangeable.

Main result · DeepSeek-7B

High apparent forget quality can mask substantial recoverability.

Across attributes, standard decoding often indicates strong forgetting. RIG or RG can sharply lower the measured forget quality, showing that target information remains accessible through the model’s logits.

Four bar charts for blood type, year of birth, postcode, and social insurance number. Original forget quality is high for many unlearning methods, while RIG or RG substantially lowers it; retraining from scratch remains comparatively stable. View full-size figure
Original forget quality (red) compared with RIG recovery (green) and RG recovery (blue) across the four FPI attributes.
Postcode · GA+KL ≈97% of supposedly forgotten postcodes are recovered under RIG.

Standard versus audited. The red bars show original forget quality; green and blue show RIG and RG audits. A lower audited score means predictions move closer to the true PI. PO is audited with RG and reported separately in the paper.

Why retraining matters. The retrained model f* never sees forget samples and does not show the same consistent audit-induced decline, providing the gold-standard reference for approximate unlearning.

Additional analyses

The recovery gap persists across models and training dynamics.

Supporting experiments reproduce the principal comparison on Qwen3-8B and trace how apparent forgetting and audited recovery diverge during GD unlearning.

Four bar charts reproducing the original and audited forget-quality comparison on Qwen3-8B across the four FPI attributes. View larger

Cross-model replication

The main original-versus-audited forget-quality comparison is reproduced on Qwen3-8B.

Four line charts showing standard forget quality generally rising while forget quality under RIG falls across GD training iterations for blood type, year of birth, postcode, and social insurance number. View larger

Unlearning dynamics

During GD training, standard forget quality rises while audited forget quality under RIG generally falls.

More analysesLearning-rate trade-offs and restricted beam-search auditing

Learning-rate trade-offs

The paper finds a recurring tension: settings that resist RIG often incur severe utility degradation, while utility-preserving settings tend to leave recoverable traces. Each plot compares GA, GD, GA+KL, and NPO after 50 unlearning iterations.

Learning-rate ablation comparing forget quality, audited recovery, and utility for blood type.
Blood type
Learning-rate ablation comparing forget quality, audited recovery, and utility for year of birth.
Year of birth
Learning-rate ablation comparing forget quality, audited recovery, and utility for postcode.
Postcode
Learning-rate ablation comparing forget quality, audited recovery, and utility for social insurance number.
Social insurance number

Restricted beam-search auditing

For NPO on postcode and SIN, restricted beam search finds lower-error candidate sets than random guessing, showing that traces can remain even when the simpler RG probe finds little.

Line chart comparing minimum error from restricted beam search with random guessing for postcode and social insurance number as the number of candidate sequences grows.
Restricted beam search compared with random guessing.

Project materials

Resources

Read the paper, inspect the official implementation, or use the FPI benchmark to reproduce and extend the study.

Citation

BibTeX

@misc{hu2026recoverability,
  title={On the Recoverability of Private Information Unlearning in Large Language Models},
  author={Shicheng Hu and Runzhi Tian and Ziqiao Wang and Yongyi Mao},
  year={2026},
  eprint={2608.29943},
  archivePrefix={arXiv},
  primaryClass={cs.LG},
  url={https://arxiv.org/abs/2608.29943}
}