Independent researcher building evidence-gated AI systems that must learn from action, feedback, and failure.
AI systems can produce fluent answers. My work asks the harder operational questions: What changes after the system acts? Does a correction survive the next interaction? What evidence must exist before a consequential action is authorized?
- AI Agent Reliability Lab - four public routes for testing observer effects, repeat failure, action evidence, and failure memory.
- KDD 2026 research feature - two non-archival workshop papers on ReflexBench and Cognitive Immunity.
- Open research community - learn, run an artifact, contribute a bounded improvement, or propose a concrete reliability problem.
- Paper and DOI index - public manuscripts, versions, records, and stated claim boundaries.
| Project | Public question | First action |
|---|---|---|
| ReflexBench | Does agent reasoning survive when its output changes users, evidence, incentives, or institutions? | Inspect 20 released scenarios and four observer-depth levels. |
| WisdomBench | Does feedback produce durable change, or does the system repeat the same failure? | Recompute the bundled longitudinal metrics. |
| Proof-Carrying Action | What evidence must accompany a consequential AI action request? | Run the no-credit repair demo and inspect the proof schema. |
| SOVEREIGN public interfaces | How can verified failure become a scoped, reviewable, reversible rule? | Run the deterministic failure-memory lifecycle fixture. |
- Ouroboros manuscript sources
- Reflexive Intelligence paper and source
- Hugging Face datasets and technical mirror
- Zenodo portfolio
Questions, safe use cases, documentation repairs, public baselines, negative results, reproducible extensions, and narrow interoperability work are welcome.
The public repositories expose papers, protocols, schemas, validators, small fixtures, minimal references, and documented limitations. They do not expose production orchestration, exact operational thresholds or weights, private prompts or data, customer systems, deployment automation, or unreleased research. A public benchmark result is not production safety certification.
Full boundary: OPEN_SOURCE_BOUNDARY.md