I'm a Data Scientist and a graduate of IIT Guwahati. I don't stick to one domain β I go wherever a hackathon, competition, or an interesting problem takes me. That's meant pathology one stretch, large language models the next, computer vision after that, and classical ML and data engineering running underneath most of it.
This GitHub is where I put personal projects I think are useful to others, not a showcase of one specialty. If there's a challenge worth taking on, I'm probably going to take it on β regardless of which domain it happens to live in.
- 𧬠Computational pathology & spatial biology β predicting molecular signals from tissue images, benchmarking pathology foundation models.
- π§ Large language models β training and fine-tuning foundation models, including a 5B-parameter model across 100+ H100 GPUs.
- π€ Multi-agent systems β agent crews that read literature, cluster findings, and surface insights.
- ποΈ Computer vision β from reading vitals off ICU monitor screens to building a vision transformer from scratch.
- π Classical ML & data engineering β clustering, topic modeling, and NLP pipelines at scale.
- π Hackathons & competitions β Inter IIT Tech Meet, Kaggle, EY Open Science Challenge, and whatever's next.
- βοΈ Infrastructure β self-hosted MLOps stacks for running experiments end to end.
None of these is "the" thing I do β they're just where curiosity and a good challenge have taken me so far.
| Project | Domain | What it does |
|---|---|---|
| Biomedical Foundation LLM | LLMs | A 5B-parameter life-sciences language model β 64K context, trained across 100+ H100 GPUs, refined through multi-stage SFT. |
| Multi-Agent Research Assistant | Agents | A crew of AI agents that read biomedical literature, cluster findings, and surface insights β built on CrewAI. |
| RAG-Based Research Chatbot | Agents Β· LLMs | Citation-backed Q&A over scientific documents with hybrid search and multilingual voice support β cut research time by 80%. |
| Cloud-Physician β ECG Vitals Extraction | Computer Vision | An ensemble of detection and OCR models that reads vitals straight off ICU monitor screens β 95% training / 80% validation accuracy. |
| Vision Transformer, From Scratch | Computer Vision | A from-scratch ViT implementation, built to actually understand attention rather than just import it. |
| GitLab MLOps Stack | Infrastructure | Self-hosted infrastructure for running and tracking AI experiments end to end. |
| Pathology Foundation Model Benchmark | Pathology | Compares Virchow, UNI, H-Optimus, and HistoEncoder on real diagnostic tasks. |
| Loki Validation Framework | Pathology | Validates and benchmarks image-to-gene prediction pipelines for spatial transcriptomics. |
Repos for the items above marked private or in progress β happy to share details if you ask.
| Languages | |
| AI & modeling | |
| Agents & LLMs | |
| Data & infrastructure |
Most of these came from saying yes to a hackathon or competition:
- π₯ Gold Medal, Inter IIT Tech Meet 2022 β SEC Filing Analyzer for SaaS companies
- π€ Research poster, "3D BiCycle GANs Meet LLMs" β selected for NVIDIA GTC 2025
- π Global Finalist, EY Open Science Data Challenge 2023
- π Kaggle Notebooks & Discussions Master