Skip to main content

Research timeline

Related research and updates

Public articles linked to the same research event.

arXiv

GrayShield overwrites Transformer weight mantissa LSBs with a Gray-code-guided low-transition sequence, cutting attacker recovery by about 50 percentage points across four model presets and two real-world malware payloads while keeping accuracy impact under 1%

The work proposes GrayShield, a lightweight, post-training, zero-data sanitization method that completely overwrites the declared mantissa least-significant-bit channel of 32-bit floating-point Transformer weights with a Gray-code-guided low-transition sequence, giving that channel zero capacity; across four Transformer model presets, two real-world malware payloads, and five attacker variants, it maintains sub-1% accuracy impact and achieves a 49.96±0.66 percentage-point Recovery Reduction against seven post-training defenses, with substantially smaller weight-distribution shift than PatternMask and Post-Training Quantization.