Detects malicious URLs in real time from the URL string alone — no page fetch, no third-party API. Trained on ~407k labelled URLs and served behind a Flask web UI with an interactive verdict gauge and per-signal breakdown.
Two feature views are combined in a single scikit-learn pipeline:
- Character n-grams (2–5, TF-IDF). Catch what counting can't: brand
lookalikes buried in subdomains (
microsoft.login-security-update.com), throwaway TLDs, and byte-level patterns of URL obfuscation. - Nine lexical signals. Interpretable structure features that can be reliably extracted from any URL without visiting it:
| Signal | Why it matters |
|---|---|
url_length |
Phishing URLs are on average much longer |
hostname_length |
Long hosts often hide a spoofed brand name |
path_length |
Deep paths are common in compromised-site phishing |
risky_word_count |
login / verify / secure / account / update… |
hyphen_count |
secure-paypal-login.com-style domains |
digit_count |
Digit-heavy URLs correlate strongly with abuse |
special_char_count |
@, ?, =, &, %, ~, + in the URL |
subdomain_count |
Subdomain depth beyond the registered domain |
has_ip |
Legitimate sites almost never serve from a bare IP |
Both views feed a logistic regression. On a held-out 20% split (81k URLs): 98.5% accuracy, 97.9% precision, 92.4% recall, 0.951 F1, 0.996 ROC AUC.
Hostname parsing strips userinfo, so google.com@evil-site.ru is scored as
evil-site.ru — the host a browser would actually visit.
A clean-looking URL on a brand-new domain carries no lexical evidence, so two optional layers add evidence from outside the string — still no LLM, no paid API:
- Domain reputation (offline).
python fetch_tranco.pydownloads the Tranco top-1M list todata/tranco.csv. If the registered domain is ranked, the verdict is softened and the rank is shown as evidence. Suffix matching is strict:mail.google.commatchesgoogle.com, butgoogle.com.evil.tkmatches nothing. A very confident model (≥0.9) keeps its verdict — popular platforms can still host phishing pages. - Deep scan (opt-in, per click). Sends only the domain name out and reports three checks side by side: does it resolve (DNS), how old is the registration (RDAP — most phishing domains are days old), and who issued the TLS certificate. The default scan stays fully offline.
An earlier version misclassified https://www.wikipedia.org as phishing with
94% confidence. The training corpus stores legitimate URLs without scheme or
www. and always with a path, so ordinary real-world input looked
out-of-distribution — a train/inference feature mismatch. The fix is a single
normalization step (normalize_url) applied identically at training and
prediction time. It also accepts defanged URLs the way analysts share them:
hxxp://evil[.]com.
pip install -r requirements.txt
The training data is data/urls_raw.csv (columns url,label with labels
good/bad), from the public dataset at
faizann24/Using-machine-learning-to-detect-malicious-URLs.
python train.py
Cleans the raw data (drops null/short rows, dedupes on the normalized URL,
maps labels to 0/1), fits the pipeline and prints accuracy, precision, recall,
F1, ROC AUC and the confusion matrix. Saves model.joblib plus metrics.json
(including per-class signal averages, which the UI uses to show how a scanned
URL compares to typical legitimate and phishing URLs).
python app.py
Open http://127.0.0.1:5000 — paste a URL (or click an example) and get a
verdict gauge, the nine extracted signals compared against class averages, and
the model's test metrics. Scans are linkable: /?url=example.com.
Verdicts are three-way on purpose: anything the model can't call cleanly (probability 0.35–0.65) is surfaced as suspicious rather than forced into a yes/no answer.
There is also a JSON endpoint:
POST /api/check
{"url": "example.com/login"}
The model only sees the URL string. A short, clean-looking phishing URL can slip through, and unusual-but-legitimate URLs can score high. Use it as a signal, not a blocklist.