I build small, reproducible evaluation artifacts for researchers and engineers reviewing agentic AI systems.
My public work focuses on:
- agent authority boundaries
- tool-use safety
- human oversight
- auditable decision evidence
- deterministic synthetic evaluations
Public repositories contain synthetic, clean-room, or educational artifacts. They do not expose private production systems, logs, prompts, schemas, or customer data.
Evidence-Gated Evaluation of Agentic AI Systems — METHODS PREPRINT / RUNNABLE / TESTED
A bounded framework for testing agent capability without expanding authority. The public release includes the manuscript, JSON Schema, 18 synthetic cases, deterministic gate logic, regression tests, frozen results, and independent-review materials.
Start with:
frozen claim + authority envelope + evidence ledger -> PASS / HOLD / FAIL / INVALID_RUN
A passing result supports only the predeclared claim within the tested environment. It grants no authority to deploy or act.
AI Governance Benchmarks — RUNNABLE / TESTED
A clean-room benchmark suite using synthetic cases, deterministic scoring, generated reports, and regression tests.
synthetic case -> scorer -> report -> tests
Start with:
- Runnable synthetic evaluation cases
- Deterministic scoring and reproducible reports
- Tests that detect boundary failures and unsupported public claims
- Explicit separation between capability evidence and operational authority
- Agent Action Audit Template —
RUNNABLE / TESTED— schema-backed synthetic action receipts, blocked-action examples, human-review metadata, and validation tests. - Human-AI Governance Lab —
RUNNABLE / TESTED— toy workflow gates for risk classification, human approval, reports, and synthetic audit receipts. - Active Inference Primer —
RESEARCH_NOTES / UTILITIES— minimal educational free-energy utilities with synthetic numerical tests and explicit limitations. - Eudaimonic Alignment —
RESEARCH_NOTES— public research notes on human flourishing, agency, and alignment/governance questions. - Quantum AI Experiments —
SANDBOX— simulator-first quantum/AI-adjacent experiments with explicit claim boundaries.
I am open to serious collaborators in AI evaluation, agent safety, governance engineering, red-teaming, and applied research.



