AI engineer building agentic systems with evidence-based eval, not vibes.
I build AI agents that are honest about what they can and can't do yet. Every project below ships with an evidence table separating verified behavior from aspirational claims, a real test suite, and an eval harness, not just a demo.
InsightBridge is an 8-agent LangGraph pipeline that turns natural-language business questions into governed, audited SQL, insights, and charts.
PipelinePilot is a sales-ops agent that qualifies leads against an ICP, drafts outreach, and routes every action through a human approval queue before touching a CRM.
DTC Support Agent is a policy-grounded customer support agent with OTP verification, shadow-mode actions, and a 94-scenario eval harness.
Every repo above documents exactly what's verified today versus what's still a mock or a roadmap item, backed by pytest suites and CI rather than inflated claims.