Skip to content
View byshivam's full-sized avatar
🎯
Learning
🎯
Learning

Block or report byshivam

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Content in all repositories owned by your account will be closed.
Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
byshivam/README.md

Hi, I'm Shivam 👋

AI Quality Engineer — I test LLMs, RAG pipelines and AI agents so they can be released with evidence, not hope.

Focused on AI in banking and financial services: evaluation, model risk and release decisions.

 

Project What it does
banking-rag-eval RAG assistant scored on a golden set with DeepEval and deterministic checks. Every CI run ends in a GO / NO-GO gate and is archived, so regressions can be traced.
hr-agent-eval Trace-level evaluation of a tool-calling agent. Its first run caught the agent booking "next Monday" on a Friday; later, an identical re-run flipped GO → NO-GO.
ai-release-readiness Turns live eval results into a model card, risk register and NIST AI RMF mapping. It showed a "passing" hallucination check rested on just 2 cases. Dashboard →
Test automation foundation  

API, data and UI suites for one fictional bank (arya-bank-sandbox), each proven against planted bugs:

 

Certified · CLLMSP · Azure AI Engineer · IBM AI Engineering · IBM AI Product Manager · ISTQB CTAL-TA

LinkedIn →

Pinned Loading

  1. banking-rag-eval banking-rag-eval Public

    Evaluation suite for a banking RAG assistant — faithfulness, hallucination & relevance scoring with CI regression gates (DeepEval · RAGAS · pytest)

    Python 1

  2. hr-agent-eval hr-agent-eval Public

    Evaluation suite for an HR AI agent — tool-use correctness, multi-step task completion, policy refusals and trace-level scoring with CI release gates

    Python 1

  3. ai-release-readiness ai-release-readiness Public

    AI release-readiness framework — model cards, risk register and NIST AI RMF mapping generated from live evaluation results, with an automated GO / NO-GO dashboard

    Python

  4. banking-data-quality banking-data-quality Public

    Data quality and reconciliation suite for synthetic Arya Bank data (Faker + DuckDB) with planted defects: SQL checks, Great Expectations, pytest, source-to-report reconciliation, daily quality scor…

    Python

  5. netbanking-ui-testing netbanking-ui-testing Public

    Playwright + TypeScript test suite for a fictional Arya Bank net banking app with planted UI bugs: Page Object Model E2E, axe-core WCAG 2.2 accessibility, visual regression, cross-browser and mobil…

    JavaScript

  6. payments-api-testing payments-api-testing Public

    Test suite for a fictional Arya Bank payments API (FastAPI + SQLite) with planted bugs: pytest API tests, idempotency checks, Pact contract tests, DB ledger reconciliation, k6 load tests, CI.

    Java