Skip to content
View dhananjaynerkar's full-sized avatar
🎯
Focusing
🎯
Focusing

Highlights

  • Pro

Block or report dhananjaynerkar

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Content in all repositories owned by your account will be closed.
Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
dhananjaynerkar/README.md

DHANANJAY NERKAR

AI/ML Engineer Candidate · RAG · LLM Applications · Applied Machine Learning

AI and ML engineering portfolio map: document AI, retrieval, models, and applications

Verified AI and ML technology stack grouped into RAG and LLM applications, applied machine learning and evaluation, and AI application engineering

Verified Technology Icons

AI / ML and application engineering
Python TensorFlow scikit-learn FastAPI Flask PostgreSQL Docker Kubernetes GitHub Actions

Supporting web, database, and mobile technologies
TypeScript React Vite Node.js JavaScript MongoDB Firebase Kotlin Java

These individual icons are a responsive visual index of technologies evidenced across the repositories; they stay centered and can wrap naturally as the README column narrows. GitHub profile READMEs cannot run media queries or JavaScript, so the icon order is fixed while the line wrapping is fluid. Spark, SQL, Pandas, NumPy, Ollama, pgvector, and retrieval methods remain represented in the text stack because the icon source does not provide an equally precise supported icon for each one.

About

I build evidence-grounded AI applications and evaluation-aware machine-learning systems with Python.

My public work shows a consistent path from documents and datasets to retrieval or model pipelines, usable application surfaces, and explicit evaluation or provenance checks.

Focus What the repositories demonstrate
Grounded GenAI Local-first RAG, document ingestion, embeddings, hybrid search, reranking, citations, provenance, and access control
Applied ML Leakage-aware preprocessing, imbalance handling, temporal validation, threshold tuning, explainability, and business-loss framing
AI application engineering FastAPI, Flask, Streamlit, PostgreSQL-backed workflows, serialized artifacts, tests, and bounded CI

Portfolio boundary: production deployment, cloud ownership, user scale, and business impact are not claimed unless the linked repository provides evidence.

GitHub · LinkedIn · Portfolio · Resume


What I Build

GROUNDED DOCUMENT AI

PDF/OCR ingestion, hybrid retrieval, citations, ACLs, and local generation.

Port Management RAG
FRAUD AND RISK ML

Leakage-aware preprocessing, imbalance handling, temporal validation, threshold tuning, and business-loss analysis.

ClaimShield · Insurerisk
DATA AND ML ENGINEERING

RDDs, DataFrames, Spark SQL, ETL, Spark ML, Docker, and Kubernetes packaging.

Spark education analytics
NLP RETRIEVAL

Structured symptom matching, dense retrieval, multilingual normalization, and safety rules.

HealthBot

Verified Technology Stack

Area Tools and methods
Programming Python · SQL
Machine Learning scikit-learn · TensorFlow/Keras · Spark ML · feature engineering · imbalanced learning · model evaluation
GenAI and RAG Ollama · embeddings · PostgreSQL full-text search · pgvector · hybrid retrieval · reciprocal rank fusion · reranking · citations · provenance
Applications and Data FastAPI · Flask · Streamlit · PostgreSQL · PySpark · Spark SQL
Quality and Reproducibility pytest · Ruff · model/artifact serialization · configuration-driven pipelines · GitHub Actions · Docker · Kubernetes packaging

Visual Engineering Pipeline

Verified AI and ML engineering layers from documents and datasets through retrieval, machine learning, APIs, applications, evaluation, and provenance

The map is a capability view of the repositories below; it does not imply that every project contains every layer or that any system is a production service.

Engineering Toolchain

Pythondocuments / tabular dataretrieval / ML / local LLMFastAPI / Flask / StreamlitPostgreSQL / pgvectortests / artifacts / evaluationDocker / Kubernetes packaging

This is a verified capability path across the portfolio, not a claim that one repository implements every step.

Flagship System

PROBLEM SURFACE

Port-management documents and governed workflows
AI PIPELINE

PDF/OCR → full-text + vector retrieval → rank fusion and reranking → local LLM → citations and ACLs
PROOF

AnyHit@5 0.89 · EvidenceCoverage@5 0.85 · 10/10 facts · 9/9 citation-valid replays

Verified Port RAG architecture from PDF and OCR ingestion through PostgreSQL full-text and vector retrieval, ranking, access control, local generation, citation validation, and corpus-bounded evaluation

Local-first prototype · RAG and workflow platform · production deployment not verified

Python, FastAPI, PostgreSQL, pgvector, Ollama, and BGE-M3.

  • PDF/OCR ingestion with page-level provenance and quarantine states.
  • Lexical plus dense retrieval with rank fusion, reranking, ACL filtering, and citation validation.
  • Corpus-bound checkpoint: AnyHit@5 0.89, EvidenceCoverage@5 0.85, 10/10 mapped facts covered, and 9/9 citation-valid replays.
  • Includes architecture, security, evaluation, workflow, operations, and bounded CI documentation.

Selected Projects

CLAIMSHIELD

ML · FRAUD DECISION SUPPORT

TensorFlow/Keras workflow with leakage-aware preprocessing, training-only imbalance handling, threshold selection, and reloadable artifacts.

ROC-AUC 0.8161 · PR-AUC 0.1829 · recall 0.8811 at threshold 0.30

Dataset redistribution rights require confirmation.
INSURERISK

ML · EVALUATION ENGINEERING

Time-aware tabular pipeline with leakage auditing, temporal features, imbalance handling, threshold tuning, business-loss analysis, artifacts, and Streamlit scoring.

Test-window checkpoint is weak; no production success or business impact is claimed.
SMART EDUCATION ANALYTICS

DATA · SPARK ML

OULAD case study covering RDDs, DataFrames, Spark SQL, ETL, Spark ML, Docker, Kubernetes manifests, and GitHub Actions checks.

Holdout AUC 0.9706 · accuracy 0.9142 · F1 0.9142 for the documented outcome proxy.
HEALTHBOT

NLP · DENSE RETRIEVAL

Educational local NLP demo with structured symptom matching, multilingual normalization, dense similarity, and safety rules.

Not a clinical system and not presented as document RAG.

Evidence-Based Capability Matrix

Capability Evidence level Repository proof
Grounded RAG and local LLM applications PRIMARY Port RAG — retrieval, local generation, citation validation, ACLs
Applied ML and evaluation design PRIMARY ClaimShield · Insurerisk — leakage controls, validation, thresholding, task metrics
AI application engineering DEMONSTRATED Port RAG, HealthBot, and Insurerisk — APIs, local apps, readiness, artifacts, tests
Data engineering and distributed ML DEMONSTRATED Spark case study — Spark ETL, Spark ML, Docker/Kubernetes packaging, CI

Engineering Practices

  • Separate implementation evidence from deployment and production-readiness claims.
  • Treat datasets, credentials, evaluation splits, and generated artifacts as explicit boundaries.
  • Prefer reproducible configuration, model contracts, provenance notes, and failure handling.
  • Use tests and CI to verify portable behavior while documenting excluded live-database or fixture-dependent checks.

Proof of Engineering

Capability Repository proof
RAG / LLM Port RAG → retrieval, citation, ACL, and runtime evaluation
Fraud ML ClaimShield → thresholded deep-learning inference and test metrics
Evaluation-aware ML Insurerisk → temporal validation, leakage notes, and business-loss framing
Data engineering Spark case study → Spark ETL, Spark ML, Docker, Kubernetes packaging, and CI

Current Focus

Strengthening clean-checkout reproducibility, corpus-versioned evaluation, citation and grounding checks, and clearly bounded local deployment for existing AI/ML systems.

No separate “Building Now” project is listed because a current project stage was not independently verified.

Activity

GitHub activity is visible in the native contribution graph on this profile. It is supporting context, not a claim about production impact, users, or scale.

Connect

Pinned Loading

  1. ai_powered_port_management_system ai_powered_port_management_system Public

    Local-first port-management AI platform with PDF/OCR ingestion, PostgreSQL full-text search, pgvector retrieval, reranking, citations, ACLs, and local Ollama generation.

    Python

  2. claimshield claimshield Public

    TensorFlow/Keras insurance-fraud decision-support workflow with leakage-aware preprocessing, imbalance handling, threshold selection, and reloadable inference artifacts.

    Jupyter Notebook

  3. insurerisk insurerisk Public

    Evaluation-aware insurance fraud and claim-risk ML pipeline with temporal validation, leakage auditing, imbalance handling, threshold tuning, business-loss analysis, and Streamlit scoring.

    Jupyter Notebook

  4. case_study_apache_spark case_study_apache_spark Public

    PySpark education-analytics case study using RDDs, DataFrames, Spark SQL, ETL feature preparation, Spark ML, Docker, and local Kubernetes packaging.

    Jupyter Notebook