Skip to content

Repository files navigation

URLGuard — Phishing Website Detection using Machine Learning

Detects malicious URLs in real time from the URL string alone — no page fetch, no third-party API. Trained on ~407k labelled URLs and served behind a Flask web UI with an interactive verdict gauge and per-signal breakdown.

Model

Two feature views are combined in a single scikit-learn pipeline:

  • Character n-grams (2–5, TF-IDF). Catch what counting can't: brand lookalikes buried in subdomains (microsoft.login-security-update.com), throwaway TLDs, and byte-level patterns of URL obfuscation.
  • Nine lexical signals. Interpretable structure features that can be reliably extracted from any URL without visiting it:
Signal Why it matters
url_length Phishing URLs are on average much longer
hostname_length Long hosts often hide a spoofed brand name
path_length Deep paths are common in compromised-site phishing
risky_word_count login / verify / secure / account / update…
hyphen_count secure-paypal-login.com-style domains
digit_count Digit-heavy URLs correlate strongly with abuse
special_char_count @, ?, =, &, %, ~, + in the URL
subdomain_count Subdomain depth beyond the registered domain
has_ip Legitimate sites almost never serve from a bare IP

Both views feed a logistic regression. On a held-out 20% split (81k URLs): 98.5% accuracy, 97.9% precision, 92.4% recall, 0.951 F1, 0.996 ROC AUC.

Hostname parsing strips userinfo, so google.com@evil-site.ru is scored as evil-site.ru — the host a browser would actually visit.

Beyond the URL string

A clean-looking URL on a brand-new domain carries no lexical evidence, so two optional layers add evidence from outside the string — still no LLM, no paid API:

  • Domain reputation (offline). python fetch_tranco.py downloads the Tranco top-1M list to data/tranco.csv. If the registered domain is ranked, the verdict is softened and the rank is shown as evidence. Suffix matching is strict: mail.google.com matches google.com, but google.com.evil.tk matches nothing. A very confident model (≥0.9) keeps its verdict — popular platforms can still host phishing pages.
  • Deep scan (opt-in, per click). Sends only the domain name out and reports three checks side by side: does it resolve (DNS), how old is the registration (RDAP — most phishing domains are days old), and who issued the TLS certificate. The default scan stays fully offline.

The normalization lesson

An earlier version misclassified https://www.wikipedia.org as phishing with 94% confidence. The training corpus stores legitimate URLs without scheme or www. and always with a path, so ordinary real-world input looked out-of-distribution — a train/inference feature mismatch. The fix is a single normalization step (normalize_url) applied identically at training and prediction time. It also accepts defanged URLs the way analysts share them: hxxp://evil[.]com.

Setup

pip install -r requirements.txt

The training data is data/urls_raw.csv (columns url,label with labels good/bad), from the public dataset at faizann24/Using-machine-learning-to-detect-malicious-URLs.

Train

python train.py

Cleans the raw data (drops null/short rows, dedupes on the normalized URL, maps labels to 0/1), fits the pipeline and prints accuracy, precision, recall, F1, ROC AUC and the confusion matrix. Saves model.joblib plus metrics.json (including per-class signal averages, which the UI uses to show how a scanned URL compares to typical legitimate and phishing URLs).

Run the web UI

python app.py

Open http://127.0.0.1:5000 — paste a URL (or click an example) and get a verdict gauge, the nine extracted signals compared against class averages, and the model's test metrics. Scans are linkable: /?url=example.com.

Verdicts are three-way on purpose: anything the model can't call cleanly (probability 0.35–0.65) is surfaced as suspicious rather than forced into a yes/no answer.

There is also a JSON endpoint:

POST /api/check
{"url": "example.com/login"}

Limitations

The model only sees the URL string. A short, clean-looking phishing URL can slip through, and unusual-but-legitimate URLs can score high. Use it as a signal, not a blocklist.

About

Phishing URL detection with scikit-learn: char n-gram TF-IDF + lexical signals, Flask UI with domain reputation and deep scan

Resources

Stars

1 star

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages