A resumable, evidence-grounded financial-research pipeline built on LangGraph with tiered frontier models.
FinSight turns a target - a company, an industry, or a macro topic - into an institutional research report in which every quantitative claim traces to a persisted, approved evidence row. It is a clean-room rebuild of an earlier, over-orchestrated system that could spend millions of tokens yet persist zero durable evidence. The design principle here is simple: the value chain works before the orchestration is clever.
It ports the proven parts of two open-source systems - the research loop and tools of MiroThinker and the analysis / report layer of FinRobot - onto a resumable LangGraph spine with a hard evidence contract.
- An institutional report (Markdown always; PDF + DOCX with the report extra) with a fixed section scaffold per profile (company / industry / macro), inline
[E:<id>]citations on every fact, a real DCF valuation + 3-year forecast + peer table, and an honest Methodology & Gaps appendix. - A durable evidence store for every run: each cited number is a persisted row you can audit, and computed figures carry the exact recipe used to derive them.
- Resumability: a crashed or interrupted run picks up from its last checkpoint - no re-running finished work, no double-counted evidence.
- Evidence is a hard postcondition.
A section is satisfied only by claim coverage: every required claim needs at least one approved evidence row.
Evidence has three kinds, each with its own approval basis -
web(source-tier),provider(a registered L1 data adapter), andanalytic(a computed figure, approved only when its full provenance DAG resolves to approved rows). A claim with no findable evidence yields an explicit gap, never a silent empty success, and an all-gapped run is a typed failure, not a hollow "done". - Computed numbers are verified, not trusted. Analytic evidence (DCF, ratios, forecasts, peer multiples) is written only by the LLM-free analyst node, straight from the deterministic analysis engine. The critic then re-derives every cited number from a self-contained recompute recipe and blocks any tampered, fabricated, or uncited figure before it reaches the report.
- Resume is free and safe.
LangGraph's SQLite checkpointer persists every state transition, and the evidence store is sharded per
(section, producer)and atomically rewritten, so resume never re-runs finished nodes or double-counts. - Models are tiered.
A frontier reasoning model drives planning / writing / critique; a cheaper thinking model runs the high-volume research loop; a cheap model handles mechanical extraction.
Routing lives in
config/models.yaml, never hard-coded, and per-tier token usage is metered.
- Python 3.12 (LangGraph does not yet support 3.14).
- uv for dependency management and running.
- API keys for a live run (a frontier/OpenAI-compatible provider, DeepSeek, Serper, FMP; Jina optional). Tests need no keys.
# 1. install (dev + report [PDF/DOCX/charts] + live extras)
uv sync --all-extras
# 2. configure credentials
cp .env.example .env # then fill in the keys below
set -a; source .env; set +a # export them - nothing loads .env automatically
# 3. run a report
uv run finsight run --target AAPL --profile companyThe run_id is printed to stderr at the start of a run - keep it to resume or inspect the run later.
uv run finsight run --target AAPL --profile company # profiles: company | industry | macro
uv run finsight run --target "US Macro" --profile macro --request "..." # optional custom brief
uv run finsight list # run ids under ./outputs
uv run finsight status --run-id <id> # node-level progress from the checkpoint
uv run finsight resume --run-id <id> # finish a crashed/interrupted runFlags: --output-root (default ./outputs), --run-id (resume/overwrite a specific run), --request (a custom research brief).
Each run writes to outputs/<run_id>/:
<slug>.md- the report (always), plus<slug>.pdf/<slug>.docxwhen thereportextra is installed.state/evidence/<section>__<producer>.jsonl- the sharded, auditable evidence store.checkpoints.sqlite- the LangGraph checkpoint that makes the run resumable.
- Credentials live in
.env(copy.env.example). Nothing in the code loads.env- export it into your shell (set -a; source .env; set +a) or useuv run --env-file .env.VLM_API_KEY/VLM_BASE_URL/VLM_MODEL_NAME- the frontier driver tier (planner / writer / critic).DS_API_KEY/DS_BASE_URL/DS_MODEL_NAME- the research + cheap tiers (DeepSeek).SERPER_API_KEY(web search),FMP_API_KEY(financial statements),JINA_API_KEY(optional, page scraping).
- Model routing lives in
config/models.yaml: three tiers (driver/research/cheap), each with a model, a per-tierreasoningflag, and its credential env-var names. Nodes ask for a tier by name and never hard-code a model.
Four layers under a static, checkpointed LangGraph graph.
The graph is static; parallelism is dynamic via Send fan-out (one branch per section), and there is at most one critic->writer revision, so a run can never loop:
planner
-> (fan out per section: researcher || analyst)
-> coverage-join
-> writer
-> critic --(pass)--> composer
--(fail, once)--> writer
- L1 Data (
finsight/data/) - typed adapters (FMP, Finnhub, SEC/EDGAR, yfinance, FRED) returning normalized pydantic models, never list-wrapped DataFrames; each splits pureparse_*from thin liveget_*, and secrets are redacted from every error. - L1 Analysis (
finsight/analysis/) - a deterministic, LLM-free finance engine (ratios, 3-year forecast, DCF valuation + sensitivity, FinRobot competitor metrics). It imports onlydata.base, so it is reusable standalone; an uncomputable value isNone(never fabricated), and every number carries its formula, inputs, and period. - L2 Tools (
finsight/tools/) - Serper web search, Jina scrape+summarize, sandboxed code execution, matplotlib charts, and a MiroThinker-style recency-context manager. Every tool returns typed, JSON-serializable output and never raises into the caller. - L3 Graph + Nodes (
finsight/graph.py,finsight/nodes/) - planner (deterministic per-profile scaffold), researcher (bounded MiroThinker ReAct loop over scraped page content), analyst (the analytic-evidence bridge), coverage-join, writer (cites from graph state), critic (pure recompute + citation check), composer. - L4 State (
finsight/state.py,finsight/evidence.py,finsight/runner.py) - typed run state, the sharded evidence store with DAG-based approval, and run-id-keyed working directories + checkpoints.
Every cited fact is a persisted EvidenceRecord. Approval is a pure function of the persisted rows - there is no mutable "approved" flag to drift:
| kind | carries | approved when |
|---|---|---|
web |
source_url + classified source_tier (+ a raw quote span) |
the tier meets the claim's minimum tier |
provider |
a registered L1 adapter name | it came from a registered data adapter (not a URL host tier) |
analytic |
a formula + input_evidence_ids + period |
every input id resolves to an already-approved row (a deterministic, cycle-guarded DAG) |
A section is satisfied only when every required claim has an approved row; otherwise it is an explicit, typed gap disclosed in the report.
finsight/ the package
data/ L1 typed data adapters
analysis/ L1 deterministic finance engine (isolated, reusable)
tools/ L2 search / scrape / code / charts / recency
nodes/ L3 planner, researcher, analyst, coverage, writer, critic
graph.py the static LangGraph spine + Send fan-out
evidence.py the sharded evidence store + DAG approval
runner.py run / resume / status, keyed by run_id
cli.py the `finsight` CLI entry point
config/models.yaml tiered model routing
tests/ hermetic unit + e2e suite
EVAL.md the acceptance gate and what each e2e test pins
uv sync --all-extras
uv run pytest -q # 171 tests, offline and hermetic
uv run ruff check .Tests are fully hermetic: a scripted LLM and fixture tools drive the real graph, evidence store, and SQLite checkpointer, so CI makes no live paid-API calls.
A live end-to-end smoke run is user-triggered.
See EVAL.md for the acceptance gate and exactly which guarantee each end-to-end test pins.
See LICENSE.