Public articles linked to the same research event.
arXiv The work presents Whiteout, a practical tool that, upon requests by individuals, prevents LLMs from regurgitating their genuine personally sensitive information (PSI) by overwriting it with precise and carefully designed obfuscation samples; evaluated on modern LLMs of varying sizes and makers, including a widely-used OpenAI model, Whiteout effectively prevents disclosure of the targeted PSIs, has negligible impact on model utility and safety, outperforms existing alternatives, and is tested against countermeasures ranging from black-box attacks such as jailbreaking to white-box adaptive attacks such as relearning and quantization.
The work presents Whiteout, a practical tool that, upon requests by individuals, prevents LLMs from regurgitating their genuine personally sensitive information (PSI) by overwriting it with precise and carefully designed obfuscation samples; evaluated on modern LLMs of varying sizes and makers, including a widely-used OpenAI model, Whiteout effectively prevents disclosure of the targeted PSIs, has negligible impact on model utility and safety, outperforms existing alternatives, and is tested against countermeasures ranging from black-box attacks such as jailbreaking to white-box adaptive attacks such as relearning and quantization.
The work presents Whiteout, a practical tool that, upon requests by individuals, prevents LLMs from regurgitating their genuine personally sensitive information (PSI) by overwriting it with precise and carefully designed obfuscation samples; evaluated on modern LLMs of varying sizes and makers, including a widely-used OpenAI model, Whiteout effectively prevents disclosure of the targeted PSIs, has negligible impact on model utility and safety, outperforms existing alternatives, and is tested against countermeasures ranging from black-box attacks such as jailbreaking to white-box adaptive attacks such as relearning and quantization.
The work presents Whiteout, a practical tool that, upon requests by individuals, prevents LLMs from regurgitating their genuine personally sensitive information (PSI) by overwriting it with precise and carefully designed obfuscation samples; evaluated on modern LLMs of varying sizes and makers, including a widely-used OpenAI model, Whiteout effectively prevents disclosure of the targeted PSIs, has negligible impact on model utility and safety, outperforms existing alternatives, and is tested against countermeasures ranging from black-box attacks such as jailbreaking to white-box adaptive attacks such as relearning and quantization.