AI / ML and application engineering
Supporting web, database, and mobile technologies
These individual icons are a responsive visual index of technologies evidenced across the repositories; they stay centered and can wrap naturally as the README column narrows. GitHub profile READMEs cannot run media queries or JavaScript, so the icon order is fixed while the line wrapping is fluid. Spark, SQL, Pandas, NumPy, Ollama, pgvector, and retrieval methods remain represented in the text stack because the icon source does not provide an equally precise supported icon for each one.
I build evidence-grounded AI applications and evaluation-aware machine-learning systems with Python.
My public work shows a consistent path from documents and datasets to retrieval or model pipelines, usable application surfaces, and explicit evaluation or provenance checks.
| Focus | What the repositories demonstrate |
|---|---|
| Grounded GenAI | Local-first RAG, document ingestion, embeddings, hybrid search, reranking, citations, provenance, and access control |
| Applied ML | Leakage-aware preprocessing, imbalance handling, temporal validation, threshold tuning, explainability, and business-loss framing |
| AI application engineering | FastAPI, Flask, Streamlit, PostgreSQL-backed workflows, serialized artifacts, tests, and bounded CI |
Portfolio boundary: production deployment, cloud ownership, user scale, and business impact are not claimed unless the linked repository provides evidence.
GitHub · LinkedIn · Portfolio · Resume
| GROUNDED DOCUMENT AI PDF/OCR ingestion, hybrid retrieval, citations, ACLs, and local generation. Port Management RAG |
FRAUD AND RISK ML Leakage-aware preprocessing, imbalance handling, temporal validation, threshold tuning, and business-loss analysis. ClaimShield · Insurerisk |
| DATA AND ML ENGINEERING RDDs, DataFrames, Spark SQL, ETL, Spark ML, Docker, and Kubernetes packaging. Spark education analytics |
NLP RETRIEVAL Structured symptom matching, dense retrieval, multilingual normalization, and safety rules. HealthBot |
| Area | Tools and methods |
|---|---|
| Programming | Python · SQL |
| Machine Learning | scikit-learn · TensorFlow/Keras · Spark ML · feature engineering · imbalanced learning · model evaluation |
| GenAI and RAG | Ollama · embeddings · PostgreSQL full-text search · pgvector · hybrid retrieval · reciprocal rank fusion · reranking · citations · provenance |
| Applications and Data | FastAPI · Flask · Streamlit · PostgreSQL · PySpark · Spark SQL |
| Quality and Reproducibility | pytest · Ruff · model/artifact serialization · configuration-driven pipelines · GitHub Actions · Docker · Kubernetes packaging |
The map is a capability view of the repositories below; it does not imply that every project contains every layer or that any system is a production service.
Python → documents / tabular data → retrieval / ML / local LLM → FastAPI / Flask / Streamlit → PostgreSQL / pgvector → tests / artifacts / evaluation → Docker / Kubernetes packaging
This is a verified capability path across the portfolio, not a claim that one repository implements every step.
| PROBLEM SURFACE Port-management documents and governed workflows |
AI PIPELINE PDF/OCR → full-text + vector retrieval → rank fusion and reranking → local LLM → citations and ACLs |
PROOF AnyHit@5 0.89 · EvidenceCoverage@5 0.85 · 10/10 facts · 9/9 citation-valid replays |
Local-first prototype · RAG and workflow platform · production deployment not verified
Python, FastAPI, PostgreSQL, pgvector, Ollama, and BGE-M3.
- PDF/OCR ingestion with page-level provenance and quarantine states.
- Lexical plus dense retrieval with rank fusion, reranking, ACL filtering, and citation validation.
- Corpus-bound checkpoint: AnyHit@5 0.89, EvidenceCoverage@5 0.85, 10/10 mapped facts covered, and 9/9 citation-valid replays.
- Includes architecture, security, evaluation, workflow, operations, and bounded CI documentation.
| CLAIMSHIELD ML · FRAUD DECISION SUPPORT TensorFlow/Keras workflow with leakage-aware preprocessing, training-only imbalance handling, threshold selection, and reloadable artifacts. ROC-AUC 0.8161 · PR-AUC 0.1829 · recall 0.8811 at threshold 0.30 Dataset redistribution rights require confirmation. |
INSURERISK ML · EVALUATION ENGINEERING Time-aware tabular pipeline with leakage auditing, temporal features, imbalance handling, threshold tuning, business-loss analysis, artifacts, and Streamlit scoring. Test-window checkpoint is weak; no production success or business impact is claimed. |
| SMART EDUCATION ANALYTICS DATA · SPARK ML OULAD case study covering RDDs, DataFrames, Spark SQL, ETL, Spark ML, Docker, Kubernetes manifests, and GitHub Actions checks. Holdout AUC 0.9706 · accuracy 0.9142 · F1 0.9142 for the documented outcome proxy. |
HEALTHBOT NLP · DENSE RETRIEVAL Educational local NLP demo with structured symptom matching, multilingual normalization, dense similarity, and safety rules. Not a clinical system and not presented as document RAG. |
| Capability | Evidence level | Repository proof |
|---|---|---|
| Grounded RAG and local LLM applications | PRIMARY | Port RAG — retrieval, local generation, citation validation, ACLs |
| Applied ML and evaluation design | PRIMARY | ClaimShield · Insurerisk — leakage controls, validation, thresholding, task metrics |
| AI application engineering | DEMONSTRATED | Port RAG, HealthBot, and Insurerisk — APIs, local apps, readiness, artifacts, tests |
| Data engineering and distributed ML | DEMONSTRATED | Spark case study — Spark ETL, Spark ML, Docker/Kubernetes packaging, CI |
- Separate implementation evidence from deployment and production-readiness claims.
- Treat datasets, credentials, evaluation splits, and generated artifacts as explicit boundaries.
- Prefer reproducible configuration, model contracts, provenance notes, and failure handling.
- Use tests and CI to verify portable behavior while documenting excluded live-database or fixture-dependent checks.
| Capability | Repository proof |
|---|---|
| RAG / LLM | Port RAG → retrieval, citation, ACL, and runtime evaluation |
| Fraud ML | ClaimShield → thresholded deep-learning inference and test metrics |
| Evaluation-aware ML | Insurerisk → temporal validation, leakage notes, and business-loss framing |
| Data engineering | Spark case study → Spark ETL, Spark ML, Docker, Kubernetes packaging, and CI |
Strengthening clean-checkout reproducibility, corpus-versioned evaluation, citation and grounding checks, and clearly bounded local deployment for existing AI/ML systems.
No separate “Building Now” project is listed because a current project stage was not independently verified.
GitHub activity is visible in the native contribution graph on this profile. It is supporting context, not a claim about production impact, users, or scale.
- Email: nerkarr.dhananjay@gmail.com
- LinkedIn: https://www.linkedin.com/in/dhananjay-nerkar
- Portfolio: https://portfolio-dhananjay-pi.vercel.app/
- GitHub Pages: https://dhananjaynerkar.github.io/
- GitHub: https://github.com/dhananjaynerkar