Live demo: https://settle-wise-ten.vercel.app · Video: 2-minute walkthrough
A decision-support system for debt collection in which models recommend and code decides. An R analytics layer - a behavioural similarity network, tested statistics, a small predictive model, and impact-network-style scenario analysis - sits over an operational system where an LLM agent negotiates with borrowers but can only act through tools whose limits are enforced deterministically. The data is synthetic with a planted structure, so every method can be scored against known truth.
It began as a hackathon voice-agent demo. The part worth reading is what grew underneath it.
- Open the Intelligence page. Everything in the analysis section below is rendered there, from the same tables the pipeline writes.
- Reproduce it - a handful of
maketargets build the synthetic history, the operational database and the whole R pipeline (see Run it locally). - Read the code -
intelligence/R/is eight numbered scripts, one stage each, every one writing its tables back to the database the Python API reads.
1,000 synthetic historical borrowers and 26,008 interaction events (call
attempts, promises, reminders, payments), generated by
server/intelligence/synthetic.py with five planted behavioural
communities. Synthetic is a deliberate choice, not a stand-in for data I
couldn't get: it means the community recovery, the effect estimates and the
model can all be checked against a ground truth that real collections data
never has. The cost of that choice is stated under limitations.
| Stage | What it does | Result |
|---|---|---|
01_features.R |
Per-borrower behaviour vector from the event log: contact rates by time of day, objection/refusal/promise rates, call volume | 1,000 x 8 |
02_network.R |
Cosine similarity on the standardised vector, each node linked to its 20 nearest, Louvain communities; a degree-preserving rewiring null model for modularity | 13,079 edges, 5 communities, modularity 0.68 against a degree-matched rewiring null of 0.17 ± 0.004, ARI vs planted truth 0.42 |
03_statistics.R |
Hypothesis tests with borrower-clustered bootstrap CIs and Benjamini-Hochberg adjustment; nothing is reported without its n and interval | Evening calls raise pick-up, OR 1.84 (1.73-1.97, n = 15,902) - but the effect is concentrated: 7.05 in one segment, ~1.0 in three others. SMS reminder before a promised payment: OR 2.27 (1.94-2.65) |
04_model.R |
Payment-within-7-days on as-of features (no future leakage; scaler fit on train only); elastic-net vs gradient boosting, champion chosen on validation PR-AUC, calibration reported | Test ROC-AUC 0.69, PR-AUC 0.43, Brier 0.17 on 2,804 held-out attempts - modest, and shown as such |
05_evaluation.R |
Scores segments and predictions against the planted truth | ARI, per-segment payment rates vs the generator |
06_epidemiology.R |
Restates the survival data as S -> Active -> Recovered/Escalated curves and an R_eff = beta x contact rate x duration | Labelled explicitly: not an epidemic - nothing transmits between borrowers; R_eff is a load indicator |
07_percolation.R |
Removes borrowers by betweenness vs at random, 30 draws, and asks when the network fragments | Robust to 65% removal by construction (min degree 20); targeted removal diverges past 0.65 |
08_scenarios.R |
Impact-network-style scenario analysis: give an intervention to k borrowers - by network position, by risk, by reachability, or at random - with segment effect sizes drawn from their CIs on each of 200 realizations | Evening calling aimed at the least likely to pay: +11.9 payers vs +2.1 random at k = 100 (6.4 SD). Targeting by betweenness is indistinguishable from random, and the panel says why |
The honest nulls are reported alongside: the Cox hazard ratios for contact strategy on time-to-payment are not significant, the overall strategy effect is Cramér's V 0.015, and weekday does nothing.
08_scenarios.R maps onto the Garrett Lab's impact network analysis
(INAscene) parameter for parameter: where the intervention starts
(initinfo) is a targeting rule and budget; whether it takes hold
(probadopt) is the borrower's own historical contact rate; the effect size
and its uncertainty (maneffmean, maneffsd) are the segment odds ratios
above, drawn from their confidence intervals so a non-significant effect
stays small and noisy rather than being switched off; realizations use
common random numbers so every rule is a paired comparison against the
random null.
What it deliberately lacks is INA's second layer. Nothing spreads between borrowers, so there is no dispersal network - and that is exactly why betweenness targeting equals random here. Position on a similarity network says who resembles whom, not who will respond.
- Synthetic data. Effect sizes are properties of the generator, not of any real population. What transfers is the method and its checks.
- ARI 0.42 is moderate. The generator plants noise; recovery is partial by design, and the degree-matched null is what shows the structure is real. (An earlier version reported that null as 0.00 ± 0.00 - Louvain at the detection resolution collapses a random graph to one community, a zero-variance null that looked like evidence and wasn't. Both nulls are now computed and the conventional one is the comparison.)
- The model is modest (AUC 0.69), and the scenario panel's baseline is the segment's observed rate rather than a per-borrower score - scoring the 1,000 borrowers the model was trained on would be in-sample.
- Single-layer network. See above.
- Odds ratios estimated on pick-up and promise completion are applied to payment odds directly in the scenarios, which overstates the effect wherever a pick-up gain does not convert.
This project was built with an AI coding assistant, and the interesting part is the validation, because the system itself embodies the same rule: an LLM may recommend, deterministic code decides.
- The agent cannot invent a number. Every figure it says comes from a
tool call;
server/offer_engine.pyreturns an empty offer list below the floor, so there is nothing for the model to say yes to, however it is pushed. The discount cap works the same way. - Two adversarial rounds (23 tool-layer attacks with the model treated as fully compromised, plus 22 scripted conversations) found real defects that tests then locked down: a $999,999 payment link that settled a $5,050 debt, a negative link that raised the balance, two live links from one renegotiation, a NaN discount approved because NaN fails every comparison, and a stored XSS where a call summary written by the model executed in the staff dashboard. All fixed; all covered by regression tests.
- What the agent got right without help is recorded too: a caller in crisis was escalated and never asked for money; a 15-year-old got nothing disclosed; cease-and-desist is enforced by code, not by the model remembering.
- The R pipeline had its own bugs from AI-assisted drafts - an xgboost API change, a calibration function mixing per-row and per-group lengths, a leakage path through neighbour features - each caught by running it and reading the output, not by trusting it.
43 Python tests and 22 R checks run in CI; ruff is clean.
Pick a borrower on the dashboard, click Call borrower, and the agent confirms who it is speaking to before disclosing anything, states the instalment due today (one cycle's worth, never the whole balance), negotiates downward but never below the floor, texts a payment link, records the outcome, and escalates to a human on a dispute, refusal or abuse. Afterwards the borrower's page shows the transcript beside what the agent actually did - every tool call in plain English.
Dashboard "Call borrower" -> POST /api/debts/{id}/call -> Vapi --PSTN--> phone
Vapi runs the voice loop and calls back for every tool:
POST /api/vapi/tools -> tool_registry -> tools.py -> the database
Repayment runs on a cycle. due_now_percent of what's left is due now
and again every cycle_days until clear, with a floor of
min_payment_today_percent underneath. Any customer can carry their own
terms. SMS goes through Twilio (server/sms_client.py, one REST call, no
SDK) and is only sent when LIVE_SMS=true; the demo never needs it.
The 21 agent tools, the prompt (server/agent/prompt.md - edit it directly)
and the deterministic simulator that lets the demo clock replay a week of
activity in one request all live in server/agent/.
make setup # venv + Python deps; R deps via renv
make synth # regenerate the synthetic history (seeded)
make seed # operational database, including the live book
make intelligence # extract events, run the whole R pipeline
make run # API + dashboard at http://127.0.0.1:8787/dashboard/
make test # 43 Python tests + 22 R checksOr talk to the agent in text, with every tool call printed:
.venv/bin/python -m server.agent.console live_0002| Path | What's in it |
|---|---|
intelligence/R/ |
the eight-stage R pipeline described above; run_all.R runs it, tests/run_tests.R checks the pure helpers |
intelligence/renv.lock |
pinned R dependencies |
server/intelligence/ |
extracts events for R, serves its tables read-only, and turns them into a per-borrower recommendation (which the offer engine may then decline) |
server/offer_engine.py |
what may be offered - the enforced limits |
server/agent/ |
prompt, the 21-tool registry, tool implementations, the call simulator |
server/routes/ |
dashboard + customer API, Vapi webhooks, SMS, mock checkout |
dashboard/ |
operator UI, plain JS, no build step |
scripts/ |
Postgres migration and intelligence sync for the deployment |
The public demo runs on Vercel with Supabase Postgres. The R layer runs
locally and its tables are pushed with scripts/sync_intelligence.py; the
same code runs on SQLite locally and on Postgres deployed, through one
adapter in server/db.py. There is no SQLite fallback on the deployment -
the entrypoint refuses to start without DATABASE_URL.
- Identity checking is verbal plus an optional last-four check, which is rate-limited; third-party disclosure rules are the main protection.
- No real money moves.
/pay/{id}is a mock checkout. - Compliance is demo-grade - modelled on real collections practice, no legal review.
- Deleting a customer is permanent, by design, to clear test entries.
- Real SMS needs a Twilio account and, beyond a trial's verified numbers, US A2P 10DLC registration - a carrier rule, not a Twilio one.