Skip to content

Latest commit

 

History

1 Commit

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

FinSight

A resumable, evidence-grounded financial-research pipeline built on LangGraph with tiered frontier models.

FinSight turns a target - a company, an industry, or a macro topic - into an institutional research report in which every quantitative claim traces to a persisted, approved evidence row. It is a clean-room rebuild of an earlier, over-orchestrated system that could spend millions of tokens yet persist zero durable evidence. The design principle here is simple: the value chain works before the orchestration is clever.

It ports the proven parts of two open-source systems - the research loop and tools of MiroThinker and the analysis / report layer of FinRobot - onto a resumable LangGraph spine with a hard evidence contract.


What you get

  • An institutional report (Markdown always; PDF + DOCX with the report extra) with a fixed section scaffold per profile (company / industry / macro), inline [E:<id>] citations on every fact, a real DCF valuation + 3-year forecast + peer table, and an honest Methodology & Gaps appendix.
  • A durable evidence store for every run: each cited number is a persisted row you can audit, and computed figures carry the exact recipe used to derive them.
  • Resumability: a crashed or interrupted run picks up from its last checkpoint - no re-running finished work, no double-counted evidence.

Core guarantees

  • Evidence is a hard postcondition. A section is satisfied only by claim coverage: every required claim needs at least one approved evidence row. Evidence has three kinds, each with its own approval basis - web (source-tier), provider (a registered L1 data adapter), and analytic (a computed figure, approved only when its full provenance DAG resolves to approved rows). A claim with no findable evidence yields an explicit gap, never a silent empty success, and an all-gapped run is a typed failure, not a hollow "done".
  • Computed numbers are verified, not trusted. Analytic evidence (DCF, ratios, forecasts, peer multiples) is written only by the LLM-free analyst node, straight from the deterministic analysis engine. The critic then re-derives every cited number from a self-contained recompute recipe and blocks any tampered, fabricated, or uncited figure before it reaches the report.
  • Resume is free and safe. LangGraph's SQLite checkpointer persists every state transition, and the evidence store is sharded per (section, producer) and atomically rewritten, so resume never re-runs finished nodes or double-counts.
  • Models are tiered. A frontier reasoning model drives planning / writing / critique; a cheaper thinking model runs the high-volume research loop; a cheap model handles mechanical extraction. Routing lives in config/models.yaml, never hard-coded, and per-tier token usage is metered.

Requirements

  • Python 3.12 (LangGraph does not yet support 3.14).
  • uv for dependency management and running.
  • API keys for a live run (a frontier/OpenAI-compatible provider, DeepSeek, Serper, FMP; Jina optional). Tests need no keys.

Quick start

# 1. install (dev + report [PDF/DOCX/charts] + live extras)
uv sync --all-extras

# 2. configure credentials
cp .env.example .env          # then fill in the keys below
set -a; source .env; set +a   # export them - nothing loads .env automatically

# 3. run a report
uv run finsight run --target AAPL --profile company

The run_id is printed to stderr at the start of a run - keep it to resume or inspect the run later.

Usage

uv run finsight run --target AAPL --profile company   # profiles: company | industry | macro
uv run finsight run --target "US Macro" --profile macro --request "..."   # optional custom brief
uv run finsight list                                  # run ids under ./outputs
uv run finsight status --run-id <id>                  # node-level progress from the checkpoint
uv run finsight resume --run-id <id>                  # finish a crashed/interrupted run

Flags: --output-root (default ./outputs), --run-id (resume/overwrite a specific run), --request (a custom research brief).

What a run produces

Each run writes to outputs/<run_id>/:

  • <slug>.md - the report (always), plus <slug>.pdf / <slug>.docx when the report extra is installed.
  • state/evidence/<section>__<producer>.jsonl - the sharded, auditable evidence store.
  • checkpoints.sqlite - the LangGraph checkpoint that makes the run resumable.

Configuration

  • Credentials live in .env (copy .env.example). Nothing in the code loads .env - export it into your shell (set -a; source .env; set +a) or use uv run --env-file .env.
    • VLM_API_KEY / VLM_BASE_URL / VLM_MODEL_NAME - the frontier driver tier (planner / writer / critic).
    • DS_API_KEY / DS_BASE_URL / DS_MODEL_NAME - the research + cheap tiers (DeepSeek).
    • SERPER_API_KEY (web search), FMP_API_KEY (financial statements), JINA_API_KEY (optional, page scraping).
  • Model routing lives in config/models.yaml: three tiers (driver / research / cheap), each with a model, a per-tier reasoning flag, and its credential env-var names. Nodes ask for a tier by name and never hard-code a model.

Architecture

Four layers under a static, checkpointed LangGraph graph. The graph is static; parallelism is dynamic via Send fan-out (one branch per section), and there is at most one critic->writer revision, so a run can never loop:

planner
   -> (fan out per section:  researcher  ||  analyst)
   -> coverage-join
   -> writer
   -> critic  --(pass)-->  composer
              --(fail, once)-->  writer
  • L1 Data (finsight/data/) - typed adapters (FMP, Finnhub, SEC/EDGAR, yfinance, FRED) returning normalized pydantic models, never list-wrapped DataFrames; each splits pure parse_* from thin live get_*, and secrets are redacted from every error.
  • L1 Analysis (finsight/analysis/) - a deterministic, LLM-free finance engine (ratios, 3-year forecast, DCF valuation + sensitivity, FinRobot competitor metrics). It imports only data.base, so it is reusable standalone; an uncomputable value is None (never fabricated), and every number carries its formula, inputs, and period.
  • L2 Tools (finsight/tools/) - Serper web search, Jina scrape+summarize, sandboxed code execution, matplotlib charts, and a MiroThinker-style recency-context manager. Every tool returns typed, JSON-serializable output and never raises into the caller.
  • L3 Graph + Nodes (finsight/graph.py, finsight/nodes/) - planner (deterministic per-profile scaffold), researcher (bounded MiroThinker ReAct loop over scraped page content), analyst (the analytic-evidence bridge), coverage-join, writer (cites from graph state), critic (pure recompute + citation check), composer.
  • L4 State (finsight/state.py, finsight/evidence.py, finsight/runner.py) - typed run state, the sharded evidence store with DAG-based approval, and run-id-keyed working directories + checkpoints.

The evidence contract

Every cited fact is a persisted EvidenceRecord. Approval is a pure function of the persisted rows - there is no mutable "approved" flag to drift:

kind carries approved when
web source_url + classified source_tier (+ a raw quote span) the tier meets the claim's minimum tier
provider a registered L1 adapter name it came from a registered data adapter (not a URL host tier)
analytic a formula + input_evidence_ids + period every input id resolves to an already-approved row (a deterministic, cycle-guarded DAG)

A section is satisfied only when every required claim has an approved row; otherwise it is an explicit, typed gap disclosed in the report.

Project layout

finsight/            the package
  data/              L1 typed data adapters
  analysis/          L1 deterministic finance engine (isolated, reusable)
  tools/             L2 search / scrape / code / charts / recency
  nodes/             L3 planner, researcher, analyst, coverage, writer, critic
  graph.py           the static LangGraph spine + Send fan-out
  evidence.py        the sharded evidence store + DAG approval
  runner.py          run / resume / status, keyed by run_id
  cli.py             the `finsight` CLI entry point
config/models.yaml   tiered model routing
tests/               hermetic unit + e2e suite
EVAL.md              the acceptance gate and what each e2e test pins

Development

uv sync --all-extras
uv run pytest -q                     # 171 tests, offline and hermetic
uv run ruff check .

Tests are fully hermetic: a scripted LLM and fixture tools drive the real graph, evidence store, and SQLite checkpointer, so CI makes no live paid-API calls. A live end-to-end smoke run is user-triggered. See EVAL.md for the acceptance gate and exactly which guarantee each end-to-end test pins.

License

See LICENSE.

About

A resumable, evidence-grounded financial-research pipeline built on LangGraph with tiered frontier models.

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages