Official code for Steering Large Language Models using Conceptors, presented at the NeurIPS 2024 MINT Workshop.
-
Updated
Mar 13, 2025 - Jupyter Notebook
Official code for Steering Large Language Models using Conceptors, presented at the NeurIPS 2024 MINT Workshop.
Official implementation of "CSKS: Continuously Steering LLMs Sensitivity to Contextual Knowledge with Proxy Models" (EMNLP 2025)
Policy-enforcing inference for open-weight LLMs. Self-hosted and OpenAI-compatible: enforces your policy at the residual-stream level, so it survives obfuscation, role-play and jailbreak wrappers. Byte-for-byte no-op on in-policy requests.
Guided undergraduate notebooks on diffusion, flow matching, deterministic sampling, and post-hoc generative-model steering, with a validated CIFAR-10 EDM bridge.
Lobopy is a lightweight PyTorch/HuggingFace library for analysing, steering/abliteration of causal language models.
Goodfire — independent third-party profile of a public API surface, by API Evangelist. Goodfire is an AI interpretability research lab building tools to understand, debug, and intentionally design neural networks by surfacing the internal features (via sparse autoencoders) that drive model behavior.
Envariant — independent third-party profile of a public API surface, by API Evangelist. Envariant is building the control layer for foundation models — an AI interpretability SDK that lets teams inspect, steer, and control model behavior. The SDK exposes a compact set of primitives: detect and causally trace behaviors like hallucinations or invaria
To associate your repository with the model-steering topic, visit your repo's landing page and select "manage topics."