// the find
RobustNLP/CipherChat
A framework to evaluate the generalization capability of safety alignment for LLMs
CipherChat is the code release for the ICLR 2024 paper showing that GPT-4/GPT-3.5 safety alignment can be bypassed by encoding prompts as ciphers (Caesar, ASCII, Unicode, Morse) instead of natural language. It's for AI safety researchers and red-teamers who need a reproducible baseline for testing alignment robustness against non-natural-language inputs, not a general-purpose tool.
Ships the actual experimental results (hundreds of .list files loadable via torch.load) alongside the code, so findings are auditable rather than just claimed. Covers multiple ciphers and multiple target models (GPT-4, GPT-3.5, Claude-2, Llama-2 7B/13B/70B), which makes it useful for cross-model comparison. Includes real ablations (shot count, demonstration toxicity, removing the system role) rather than just the headline result, so you can see what actually drives the effect.
This is a paper artifact, not a maintained library: one main.py entry point, no requirements.txt or environment spec visible in the tree, no tests, no CI. The repo is 90% raw result dumps organized by ad-hoc filename convention, which makes it painful to extend to a new model or cipher without reading main.py end to end. No abstraction for plugging in an arbitrary target LLM or cipher — you're editing the script, not calling an API. Given it targets GPT-4-0613 and older Llama-2 checkpoints, most of the pinned results are already stale against current model safety training.