Early interaction dynamics add predictive signal beyond profile similarity and improve identification of rare durable ties. In the held-out online cohort, adding first-month dynamics raised ROC-AUC from 0.662 to 0.707 and concentrated durable ties at 11.2× the base rate in the combined model's top 1%.
Open the zero-install IRL research bridge →
The study asks whether the way two people interact during their first month predicts whether they will still interact months later—beyond what their profiles have in common. Adding first-month dynamics raised ROC-AUC from 0.662 to 0.707 on a strict chronological test set: +4.48 points, with a paired-bootstrap 95% interval of +1.93 to +7.14 points. Average precision rose from 0.021 to 0.055 (+0.033; 95% interval +0.015 to +0.065) against a 1.06% test-set prevalence.
This supports the precise claim that dynamics add useful information beyond similarity. It does not show that dynamics alone dominate every similarity signal: the dynamics-only model scored 0.642 ROC-AUC, below the 0.662 similarity model, while the combined model performed best.
| Feature set | ROC-AUC | PR-AUC | MRR | Recall@10 |
|---|---|---|---|---|
| Profile similarity | 0.662 | 0.021 | 0.311 | 0.523 |
| Interaction dynamics | 0.642 | 0.039 | 0.367 | 0.547 |
| Combined | 0.707 | 0.055 | 0.392 | 0.576 |
The online benchmark is evidence for a research direction—not a model that should be applied to people after an in-person event. A prospective IRL study would translate the design as follows:
| Research role | Group-experience data |
|---|---|
| Similarity baseline | Profile questionnaires completed before an event |
| Interaction dynamics | Directed participant-to-participant feedback after the event |
| Early investment | Follow-up initiation, response, and another accepted experience at days 7–30 |
| Durable-tie outcome | Mutually reported continued interaction during days 90–180 |
The decisive test remains the same: compare similarity-only, dynamics-only, and combined models on future events, then report the incremental lift from reciprocal experience signals. The interactive research bridge shows how two directional questionnaire responses become dyadic features without producing a compatibility score or pretending to make an IRL prediction. The IRL group-study memo specifies the data contract, product metrics, evaluation, and a two-week pilot.
Ranking makes the rare-event result concrete. Among the 19,291 held-out dyads, the combined model's top 1% contained 23 durable ties in 193 predictions: 11.9% precision and 11.2× the test-set base rate. Its top 5% contained 50 of the 205 durable ties (24.4% recall, 4.9× lift), while its top 10% contained 70 (34.1% recall, 3.4× lift).
Lift measures concentration, not certainty. Even in the top 1%, most dyads do not satisfy the days-90-to-180 reciprocal-interaction label. The model is useful for ranking a scarce outcome, not for declaring that a relationship will last.
Held-out SHAP values show what the combined model uses. Pre-contact topic overlap, pre-contact activity balance, and interaction recency have the largest average absolute contributions. The direction is not uniformly intuitive or linear, which is another reason to avoid turning individual features into relationship advice.
SHAP explains this model's predictions; it does not establish that a feature causes a relationship to persist.
The headline cohort contains 128,604 eligible r/ApplyingToCollege dyads, including 2,369 durable
ties. The final test period contains 19,291 dyads and 205 positives. Aggregate results, split counts,
ranking metrics, bootstrap intervals, and held-out SHAP values are in
artifacts/results.json.
The local Streamlit demo is the technical companion to the published Reddit benchmark. It converts objective first-month observations into the same nine features used by the combined model. It reports a model-score percentile and the observed outcome rate for that held-out percentile band—never the class-weighted XGBoost output as a personal probability. It also shows signed local SHAP contributions. Inputs stay on the local machine and are not logged or sent to an external service.
python -m pip install -e ".[demo]"
streamlit run demo/app.pyThe accompanying candidate questionnaire separates behavioral proxies represented in the current model from exploratory emotional-safety, responsiveness, vulnerability, investment, authenticity, and growth questions. The questionnaire is not scored, and its responses are not valid substitutes for the logged model features.
- Unit: an unordered pair of distinct, non-deleted Reddit authors after their first direct reply.
- Profile window: days -30 to 0 before first contact.
- Dynamics window: days 0 to 30 after first contact.
- Positive outcome: at least one direct reply in each direction during days 90 to 180.
- Negative outcome: no direct reply in days 90 to 180 while both authors remain active.
- Censored: incomplete follow-up, either author inactive, or one-directional outcome interaction.
- Split: earliest 70% train, next 15% validation, final 15% test by first-contact time.
- Primary metric: average precision; ROC-AUC and partner-ranking metrics are also reported.
The outcome is durable online interaction, not friendship, intimacy, mental health, or a causal effect of communication behavior.
| Family | Examples | Availability |
|---|---|---|
| Profile similarity | pre-contact topic overlap, activity match, shared communities | Before contact |
| Interaction dynamics | reply count, reciprocity, latency symmetry, effort balance, regularity, recency | First 30 days |
| Behavioral annotation | disclosure depth, disclosure reciprocity, supportive responses | First 30 days |
| Graph | prior-year node2vec proximity and coverage | Before each target year |
Held-out mean absolute SHAP values place lexical topic overlap, pre-contact activity balance, and interaction recency at the top of the combined model. The annotation extension is intentionally not in the headline: 216,456 outcome-blinded messages are prepared, but no provider credential was available for a real scoring run. No claim about disclosure being the strongest signal is made yet.
The corrected node2vec analysis uses annual snapshots containing only earlier-year edges. Coverage is 49.7% for 2018 dyads; graph features do not improve the combined model. That negative ablation is kept because it is more informative than forcing a graph model into the headline. A temporal GNN remains future work only after a larger graph has adequate historical coverage.
The reader consumes official ConvoKit by-subreddit archives. Raw archives are downloaded locally and
never committed. The headline uses r/ApplyingToCollege; r/Cornell is a smaller parser and cohort
feasibility check. r/ChangeMyView is supported by the interface but excluded from the current result
because its 1.27 GB compressed archive requires a scale-out runtime not available on this workstation.
python -m venv .venv
.venv\Scripts\python -m pip install -e ".[model,demo,dev]"
$env:CONNECTION_DYNAMICS_HASH_KEY = '<private-random-key>'
connection-dynamics build-panel `
data/raw/ApplyingToCollege.corpus.zip `
--output data/processed/applying_to_college_panel.csv `
--summary artifacts/private/panel_summary.json
connection-dynamics benchmark `
data/processed/applying_to_college_panel.csv `
--output artifacts/results.json `
--predictions artifacts/private/predictions.csv `
--hero-chart artifacts/hero.png `
--model-artifact artifacts/combined-model.json `
--demo-reference artifacts/demo-reference.json `
--lift-chart artifacts/rare-tie-lift.png `
--shap-chart artifacts/shap-summary.png `
--study-label 'r/ApplyingToCollege · 128,604 eligible dyads' `
--bootstrap 2000
pytest
ruff check src testsThe optional annotation workflow is outcome-blinded, resumable, and compatible with OpenAI or an
OpenAI-compatible base URL. It refuses to add disclosure features to a confirmatory benchmark unless
every dyad has complete annotation coverage. See the
ANNOTATION_PROTOCOL.md before running it.
- Leakage audit
- Data and privacy
- Annotation protocol
- Candidate questionnaire
- IRL group-study memo
- Contributing
The public repository contains code, synthetic fixtures, aggregate metrics, charts, an identifier-free tree model, and an aggregate demo reference. It does not publish usernames, raw comments, annotation manifests, panels, or row-level predictions.


