AI Quality Engineer — I test LLMs, RAG pipelines and AI agents so they can be released with evidence, not hope.
Focused on AI in banking and financial services: evaluation, model risk and release decisions.
| Project | What it does |
|---|---|
| banking-rag-eval | RAG assistant scored on a golden set with DeepEval and deterministic checks. Every CI run ends in a GO / NO-GO gate and is archived, so regressions can be traced. |
| hr-agent-eval | Trace-level evaluation of a tool-calling agent. Its first run caught the agent booking "next Monday" on a Friday; later, an identical re-run flipped GO → NO-GO. |
| ai-release-readiness | Turns live eval results into a model card, risk register and NIST AI RMF mapping. It showed a "passing" hallucination check rested on just 2 cases. Dashboard → |
Test automation foundation
API, data and UI suites for one fictional bank (arya-bank-sandbox), each proven against planted bugs:
- payments-api-testing — Java, RestAssured, Cucumber, Pact, k6
- banking-data-quality — Python, SQL, Great Expectations
- netbanking-ui-testing — Playwright, TypeScript, accessibility checks
Certified · CLLMSP · Azure AI Engineer · IBM AI Engineering · IBM AI Product Manager · ISTQB CTAL-TA

