[ICLR 2025] General-purpose activation steering library
-
Updated
Sep 18, 2025 - Python
[ICLR 2025] General-purpose activation steering library
Benchmark evaluation code for "SORRY-Bench: Systematically Evaluating Large Language Model Safety Refusal" (ICLR 2025)
We study whether categorical refusal tokens enable controllable and interpretable safety behavior in language models.
Reproducible, evergreen benchmark for LLM refusal on biological research prompts — 19 models, 141 prompts, 13,389 adjudicated trials
🔓 Ablate — directional ablation (abliteration) toolkit for open-source LLMs. Automatic censorship/refusal removal via residual-stream direction ablation, with KL-guided search, an LLM-judge harness, and one-call push to the Hub. pip install ablate-llm
CLI for keeping Claude Fable 5 prompts, skills, and API traffic in shape — lint anti-patterns, canary silent degradation, aggregate refusal analytics. Every rule cites Anthropic docs.
RAG with verifiable citations and measured refusal — retrieval scored separately (TF-IDF beats embeddings here), citations validated against chunks actually retrieved.
Public Driftmap harness: public-safe CSV suites + rubrics + run logs for drift detection, refusal integrity, injection resistance, and uncertainty tracking.
A probe suite that measures which conversation states an LLM cannot leave. Three arms, because two cannot tell obedience from token statistics; a null only counts when the design had the power to see the effect.
中文企业公开报告 Hybrid RAG:Docling 解析 · Qdrant 稠密/稀疏检索 · 查询理解硬过滤 · 带引用生成与拒答 · 文档生命周期与评测看板
Refusal verification surface for Riverbraid fail closed policy boundaries.
RefusalScope
Unified CLI/TUI to abliterate any (V)LLM (Heretic / OBLITERATUS / ErisForge) and benchmark the methods on one schema — G/P/S composite + Pareto front. Pure-stdlib core, no GPU for the core.
Locating and editing refusal in the J-space workspace with the Jacobian lens: refusal is legible ~10 layers before the first token, and only ~1/3 lives in the verbalizable workspace.
An open reproduction of feature-level activation steering with the prompt set released, showing the capability tax that behavioural metrics miss
How does this function's cost scale? Counted, not timed — and UNDETERMINED when no complexity class settles.
Training-time defense that redistributes LLM refusal via mean/covariance matching + KD, raising linear-ablation attack rank from K=1 to K≥16 (Llama-3.2-1B-Instruct)
Mechanistic interpretability of refusal behavior in Qwen2.5 models: sparse feature interventions, residual steering, judge-vs-rule analysis, and 1.5B replication
Bank and card statements into a reconciled ledger, with a plain-English refusal for the ones that do not add up: silently wrong money $62,832.40 to $0.00.
Public reference interfaces for proof-gated AI action, refusal, authority, and evidence boundaries.
To associate your repository with the refusal topic, visit your repo's landing page and select "manage topics."