閉じる

NEWS

Paper accepted at ACL 2026: “When Benchmarks Leak”

2026.07.06

Benchmark-based evaluation is the standard way to compare large language models, but its reliability is threatened by test set contamination, where test samples leak into training data and inflate reported performance.

This work proposes DeconIEP, a decontamination framework that operates entirely at evaluation time, with no retraining required. By applying small, bounded perturbations in the input embedding space, it removes the effect of contamination so that a model’s true performance can be measured.

[Paper]
Jianzhe Chai, Zhe Yu, Jun Sakuma, “When Benchmarks Leak: Inference-Time Decontamination for Large Language Models,” Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (ACL 2026), pp. 44743-44760, July 2026.

ACL 2026 website