A local research workspace for scientific literature, evidence tables, and answers traceable to source passages.
Quickstart · How it works · Architecture · Commands · Scope and limits
DEIXIS starts with a research question. It searches selected scholarly providers or accepts attached PDFs, lets the user screen and include sources, exposes passages for inspection, and produces a source-linked answer. An evidence table can extract values from selected source versions, with quotes and revision history. Research, sources, answers, and tables remain in a local SQLite library that can be reopened and backed up.
This is an AI-assisted research workspace, not an autonomous discovery system or a scientific-validity certificate. A structurally valid answer has passed deterministic schema and citation-link checks; those checks do not prove that each quoted passage supports the claim made about it.
- Ask a question and choose the academic sources and model connection for the research.
- Inspect retrieved records or add PDFs; keep or exclude sources yourself, even when model screening differs.
- Open the underlying passages and PDF pages. Reading depth, source versions, and unavailable text stay visible.
- Generate an answer with claim-to-passage links, or fill an evidence table from selected sources. Unresolved citation failures remain unverified drafts.
- Reopen the research, revise the selection, and retain earlier answers and cell revisions instead of silently replacing them.
The current checkout includes OpenAlex, Semantic Scholar, Crossref, arXiv, bioRxiv (via OpenAlex), PubMed, IEEE Xplore, Scopus, CORE, and a supplementary SerpApi search connector. Some require keys or access entitlements. Model adapters include Codex, Claude Code, Gemini, and DeepSeek; availability depends on local authentication or configured credentials, not merely on an adapter being present. Zotero collection import, BibTeX/RIS export, PDF extraction, optional OCR/equation reading, and library backup/restore are also implemented. See provider settings and decisions for conditions and boundaries.
For development, install Python 3.12, uv, Node.js/npm, and a model connection you can authenticate. The commands below build and run the local web app; they do not configure a scholarly provider or model account for you.
uv sync
cd apps/web
npm ci
npm run build
cd ../..
PYTHONPATH=backend uv run python -m deixis serveOpen http://127.0.0.1:8765/. DEIXIS listens on loopback by default; if port 8765 already serves a library, open that instance instead of starting another. On macOS the default data directory is ~/Library/Application Support/DEIXIS; set DEIXIS_DATA_DIR to an explicit separate directory for isolated work. The database, PDFs, provider payloads, and model home do not belong in Git. External provider searches and remote model calls are not offline operations.
Configure keys through the app's Connections settings or an untracked .env using the variable names in the example. Codex requires sign-in to DEIXIS's separate Codex home (by default <data directory>/codex-home); the app shows connection availability. An optional equation reader downloads its own models and is not required to start DEIXIS.
| Layer | In this repository |
|---|---|
| Local API, workflow, persistence, provider and model adapters | backend/deixis/ |
| React interface, served by the local backend after build | apps/web/ |
| Versioned model-step JSON Schemas and method instructions | contracts/research/ · methods/deixis-research/ |
| Deterministic/integration tests and synthetic browser fixtures | tests/ · apps/web/e2e/ |
| Isolated evaluation utilities, not product features | scripts/ |
The application checks source identities, reading depth, passage IDs, and exact source-owned citation anchors before publishing a linked answer. It does not infer semantic support from a valid anchor. See the documentation map, repository layout, and durable decisions for the authority of design records versus implemented code.
From the repository root, unless a command changes directory:
PYTHONPATH=backend uv run python -m pytest -q
(cd apps/web && npm run build && npm run lint)
(cd apps/web && DEIXIS_ACCEPTANCE_DIR=/tmp/deixis-acceptance npm run test:acceptance)
PYTHONPATH=backend uv run python -m deixis backup /path/to/backup-parent
DEIXIS_DATA_DIR=/path/to/empty-data-dir PYTHONPATH=backend uv run python -m deixis restore /path/to/backup-folderThe browser suite needs Google Chrome. It uses synthetic records, fixed PDFs, and a scripted model; passing it verifies application behavior, not scientific accuracy or live-provider quality. Backups include a consistent database snapshot, referenced PDFs, provider payloads, and a SHA-256 manifest, but not model credentials. Restore checks hashes and refuses a target with an existing library. Do not use the active data directory for a probe or restore.
- Search results are bounded by provider access, query choices, and budgets. Found, included, inspected, given-to-model, and cited sources are different counts; a missing result is not evidence of novelty.
- Abstract-only evidence cannot establish full-text methods, equations, or results. PDF extraction and optional OCR/equation reading have their own failure and uncertainty states.
- P4 and P5 were closed on bounded evaluations. Their reviewer judgments were delegated to a model, not independently checked by a human; known-work recall and PDF-page use remained limitations. Do not read these gates as general performance claims.
- The report, Chain of Ideas, candidate questions, and claim-specific kill-search appear in design records. Do not treat a design or method proposal as an integrated runtime feature.
DEIXIS is licensed under AGPL-3.0-or-later. It uses AGPL-licensed PyMuPDF for PDF text extraction. No CI result, published package, or hosted live site is represented by the badges above.
The material below was transferred from the earlier Quaestio design work. It preserves methodological sources, inspected-system boundaries, and accepted research-method decisions; it is not a feature list or a report of DEIXIS runtime results. For current implementation status, inspect the code and decisions.
On 14 September 2026, design documents, the prototype and local reference evidence
were moved from the adjacent Quaestio repository into this directory. The product
bibliography and design decisions below were transferred from its README.
Historical Quaestio names and revision identifiers in the records are retained.
Quaestio remains the separate methodological skill: links to ../quaestio/ require
that sibling checkout. They do not indicate an implemented integration.
Private PDFs, raw reports and screenshots remain local and ignored by Git.
Recorded on 14 September 2026 from the Luna Literature method check review. These references inform the proposed product's AI-assisted literature discovery, synthesis and hypothesis development. They support individual methodological components; they do not validate Quaestio's combined workflow, AI performance, or ability to establish scientific originality.
| Reference | Relevant contribution | Boundary when used in Quaestio |
|---|---|---|
| Levac, D., Colquhoun, H., and O'Brien, K. K. (2010). Scoping studies: advancing the methodology. Implementation Science, 5, 69. | Connect the purpose to the research question; iteratively select and chart studies; synthesize themes and refine the scope. | Health-research methodology. A conversational exploration is not automatically a formal scoping review; evidence quality also matters when identifying gaps. |
| Wohlin, C. (2014). Guidelines for snowballing in systematic literature studies and a replication in software engineering. EASE '14, article 38, pp. 1–10. Author manuscript. | Search reference lists backward and citing publications forward from relevant seed studies, iterating as needed. | A software-engineering replication supports the technique, not universal completeness. Using it alongside database queries is our workflow choice. |
| Robinson, K. A., et al. (2013). Framework for Determining Research Gaps During Systematic Review: Evaluation. AHRQ Methods Research Report, no. 13-EHC019-EF. | Describe where evidence cannot support a conclusion and why: insufficient/imprecise, biased, inconsistent, or not the right information; characterize the gap using PICOS. | Developed for health systematic reviews. Engineering comparison dimensions require adaptation; neither an empty search nor a future-work statement establishes novelty. |
| Grant, M. J., and Booth, A. (2009). A typology of reviews: an analysis of 14 review types and associated methodologies. Health Information & Libraries Journal, 26(2), 91–108. | Distinguish review approaches through Search, Appraisal, Synthesis and Analysis (SALSA), including their strengths and limitations. | A descriptive typology, not a single prescribed protocol or evidence that one AI workflow is superior. |
| Nyanchoka, L., et al. (2019). A scoping review describes methods used to identify, prioritize and display gaps in health research. Journal of Clinical Epidemiology, 109, 99–110. PubMed. | Surveys methods for identifying, prioritizing and displaying research gaps; reports the lack of standard methods at the time of the review. | A historical health-research finding, not proof that no standards exist today or that arbitrary gap scoring is valid. |
The proposed source groups (foundational work, reviews, and close recent studies), user checkpoints, candidate cards and claim-specific kill-search are product adaptations. Their usefulness and failure modes need separate evaluation. Record search scope and access limits, distinguish author statements from inference, and report an unsuccessful overlap search as no match found within the searched scope, not as a novelty verdict. See the product evidence contract and the existing skill evidence register.
Collected during the 14 September 2026 Luna repository review and Claude Opus 5
(medium) literature search. PaperLens, AI-Scientist and AI-Researcher received
bounded source-code inspection. CoI-Agent subsequently received a bounded code
review at ac94317 on the same date; the other entries are paper or README
references. Code inspection is not runtime verification. None was installed or benchmarked in Quaestio.
These references complement the methodological sources above and the extended
method comparison below.
| Reference and source links | Relevance to Quaestio | Inspection boundary |
|---|---|---|
| Chain of Ideas, Li et al. (2024), Chain of Ideas: Revolutionizing Research Via Novel Idea Development with LLM Agents. Paper, v5; official CoI-Agent repository. | Selected methodological basis for organizing literature into research-development chains and using that synthesis to develop candidate questions. | Bounded code review at ac94317: chain construction, novelty loop, prompts and selected search/text functions. Runtime behavior, retrieval coverage and reported evaluation results have not been reproduced. |
| Scideator, Radensky et al., Human-LLM Compound System for Scientific Ideation through Facet Recombination and Novelty Evaluation. Paper, v7, 2026 revision of the 2024 preprint. | Purpose/mechanism/evaluation facets, user-guided recombination, distance-controlled retrieval and facet-based comparison with prior work. | A complementary design reference, not a selected replacement for CoI. No verified code repository recorded; novelty classification is not conclusive proof of originality. |
| ResearchAgent, Baek et al. (2025), ResearchAgent: Iterative Research Idea Generation over Scientific Literature with Large Language Models. NAACL paper; PDF; preprint; official repository. | Start from a core paper, retrieve related publications and concepts, and iteratively develop problems, methods and experiment designs. | Paper and repository identity checked; code not audited. Reviewing-agent feedback does not replace original-source verification. |
| SciMON, Wang et al. (2024), SciMON: Scientific Inspiration Machines Optimized for Novelty. ACL paper; preprint; code/resource link supplied by the paper. | Literature-grounded inspiration retrieval and iterative comparison of candidate ideas with prior papers. | Paper-level reference; code not audited. Rewriting an idea to reduce similarity does not establish a substantive contribution. |
| AI-Researcher, Tang et al. (2025), AI-Researcher: Autonomous Scientific Innovation. Paper; PDF; repository; inspected survey handoff. | Structured handoffs between concepts, paper analysis, code analysis and research planning. | Bounded inspection at f9a6f84; selected entry points and memory/survey code. Benchmark/ML implementation workflows do not establish general question-first discovery performance. |
| The AI Scientist, Lu et al. (2024), The AI Scientist: Towards Fully Automated Open-Ended Scientific Discovery. Paper; repository; inspected idea/novelty code. | Separate idea generation, literature comparison, experiments and reporting. | Bounded inspection at 1de1dbc. The abstract-based search and binary novel outcome are insufficient for Quaestio's claim-specific kill-search. |
| PaperLens. Repository; inspected AI analysis; citation graph. | Natural-language paper discovery, comparison and citation-graph navigation. | Bounded inspection at 736af5a; selected search/analysis/graph routes. The inspected analysis prompt uses abstracts; a complete claim-to-passage kill-search was not established. No associated paper is asserted here. |
| Human evaluation of AI research ideas, Si, Yang and Hashimoto (2024), Can LLMs Generate Novel Research Ideas? A Large-Scale Human Study with 100+ NLP Researchers. Paper. | Evaluation design and the distinction between perceived novelty, feasibility and eventual research outcomes. | Abstract-level check. Quantitative results were not independently audited; this is not evidence of Quaestio's effectiveness. |
The inspected novelty check
returns True when no papers are retrieved and treats a parsed similarity value
other than "1" as novel. The comparison prompt
uses titles and abstracts. Its idea-regeneration loop
has no explicit attempt cap in that loop. These are code observations, not executed
failure demonstrations. Quaestio's adaptation needs explicit unknown/error states,
validated output schemas, bounded retries and claim-specific original-source
checking; it must not inherit these automatic novelty decisions.
User decision, 14 September 2026: use Chain of Ideas for Quaestio's literature synthesis and candidate-question development. The selection adopts the method; it does not mean the upstream software has been integrated or tested.
Accepted design sequence: construct the field map through research-development trajectories; make candidates concrete through purpose, mechanism and evaluation; then apply kill-search to each candidate's specific claim. CoI provides the development-chain basis; the facet structure draws on Scideator, and claim-specific adversarial checking follows Quaestio's evidence requirements.
Quaestio will adapt the method to its question-first workflow:
- Interpret the question and retrieve foundational work, relevant reviews and close recent primary studies, explaining selection beyond citation counts.
- Build evidence-linked development chains: what problem each study addressed, what it established or changed, and which uncertainty the next study pursued. Preserve competing branches and cross-links; chronology or a citation alone does not demonstrate an intellectual dependency.
- Synthesize the remaining uncertainties and plausible research directions from these chains. Preserve source versions, passages and reading-depth limits.
- Let the user steer the research direction, then describe each candidate's purpose (question or outcome), mechanism (proposed explanation or method), and evaluation (comparator and evidence needed to assess the claim). Make its claim specific enough to compare with prior work. Scale chains and candidates to the question, coverage and budget rather than enforcing fixed counts.
- Apply a distinct kill-search to each candidate's specific claim: inspect the strongest prior work in the field, methodological neighbors and adjacent applications. Report established overlap, partial overlap, no match within scope, or insufficient evidence; an empty search is not a novelty verdict.
- Present surviving or narrowed questions with conditional hypotheses, assumptions and an informative validation plan. Revise chains and claims when new evidence changes them; experiments remain a separately scoped activity.
This is Quaestio's adaptation of CoI, including its own provenance and kill-search requirements, not a verbatim reproduction of the paper. The existing modular skill remains the source-audit foundation. See the continuation brief for the accepted product decision.
The detailed research-methods document records the focused 14 September 2026 review: established methods and books, AI systems, existing repository coverage, proposed additions, guardrail ownership, candidate/claim evidence structures, future evaluation and search/access limits. It distinguishes accepted choices from recommendations and implemented behavior. This is a bounded design review, not an exhaustive literature review or a measured ranking of the most popular methods. Book inspection was limited to publisher descriptions and contents; source-specific reading levels are in that document.
Recommendation: retain the selected CoI → purpose/mechanism/evaluation → claim-specific kill-search design. Make concept comparison, assumption analysis, conditional cross-literature bridges, competing explanations and informative checks explicit inside the workflow. Add a claim-element-to-passage matrix to kill-search. These are internal methods, not additional user-facing skills. The suggested four internal responsibilities remain an architecture proposal; this documentation does not implement CoI or scientific decision gates.
Latest architecture constraint: the user does not want a multi-agent system. The first version should complete research workflows with one main agent using method files and tools. An additional bounded review is optional and separately user-requested; no standing generator/reviewer pair is required. Deterministic record, retrieval-state and budget checks are application functions. Keep brief feasibility assessment inside candidate development, begin with a concise field map before focused synthesis, and reserve human decisions for meaningful direction, scope, experiment or cost changes. Revision may end in closure or deferral. Structured PDF extraction does not by itself establish mathematical understanding.
The user also proposed an optional model-selected review: choose a candidate or report, select an available model and receive a separate evidence-linked assessment of a versioned snapshot. Findings may be fed back into the main work by the user; the reviewer does not automatically edit it. Existing evidence is the default, with extra retrieval explicitly selected. See the review design for scope, provenance and budget boundaries. A browser-only UI prototype once demonstrated model/focus selection, a separate sample report and staged feedback; it was removed (D10). Real model execution, evidence review and production persistence are not implemented.
The identifiers below match the detailed document. Existing references above and the E1–E12 skill evidence register are retained; different versions of the same work are not independent evidence.
| ID | Reference | Contribution and boundary |
|---|---|---|
| M1 | Booth, W. C., Colomb, G. G., Williams, J. M., Bizup, J., and FitzGerald, W. T. (2024). The Craft of Research, fifth edition. University of Chicago Press. | Question, problem and significance framing; publisher/contents inspection, not whole-book review or efficacy evidence. |
| M2 | Webster, J., and Watson, R. T. (2002). Analyzing the Past to Prepare for the Future: Writing a Literature Review. MIS Quarterly, 26(2), xiii–xxiii. | Concept-centric comparison; already represented in E5. Current metadata check and dated prior reading record. |
| M3 | Alvesson, M., and Sandberg, J. (2011). Generating Research Questions Through Problematization. Academy of Management Review, 36(2), 247–271. Original PDF. | Challenge consequential assumptions; conceptual methodology, adapted beyond management theory. |
| M4 | Alvesson, M., and Sandberg, J. (2024). Constructing Research Questions: Doing Interesting Research, second edition. SAGE. | Book treatment of problematization; publisher/contents inspection. Same method family as M3. |
| M5 | Swanson, D. R., and Smalheiser, N. R. (1996). Undiscovered Public Knowledge: a Ten-Year Update. KDD, 295–298. | Connections between separate literatures suggest hypotheses; they do not prove the inferred relationship. |
| M6 | Gentner, D. (1983). Structure-Mapping: A Theoretical Framework for Analogy. Cognitive Science, 7(2), 155–170. Author-hosted PDF. | Relational mapping rather than surface resemblance; transfer still needs domain validation. |
| M7 | Ritchey, T. (2018). General morphological analysis as a basic scientific modelling method. Technological Forecasting & Social Change, 126, 81–91. | Optional combination-space and consistency analysis; not a new default stage or novelty test. |
| M8 | Platt, J. R. (1964). Strong Inference. Science, 146(3642), 347–353. Original article scan. | Competing hypotheses and discriminating tests; historical methodological argument, not universal efficacy evidence. |
| M9 | Shaw, M. (2003). Writing Good Software Engineering Research Papers. ICSE, 726–736. | Match evidence to the contribution; retained E11 reading record, not a new full-text inspection. |
| M10 | Gottweis, J., et al. (2026 revision of the 2025 preprint). Accelerating scientific discovery with Co-Scientist, v2. Linked Nature publication. | Iterative hypothesis generation/critique; abstract and official overview checked. Biomedical validation is not a cross-domain guarantee. No code audit. |
| M11 | Liu, H., Zhou, Y., Li, M., Yuan, C., and Tan, C. (2025 revision). Literature Meets Data: A Synergistic Approach to Hypothesis Generation, v3. HypoGeniC/HypoRefine repository. | Optional future data-supported mode; abstract/README review, no code audit. |
| M12 | Liu, H., Huang, S., Hu, J., Zhou, Y., and Tan, C. (2026 revision). HypoBench: Towards Systematic and Principled Benchmarking for Hypothesis Generation, v2. | Real and synthetic hypothesis-generation tasks; abstract-level evaluation reference, not a benchmark already run here. |
| M13 | Si, C., Hashimoto, T., and Yang, D. (2025). The Ideation–Execution Gap: Execution Outcomes of LLM-Generated versus Human Research Ideas. | Separate proposal ratings from execution outcomes; complements the earlier Si study and E12. |
| M14 | Ikoma, H., and Mitamura, T. (2025). Can AI Examine Novelty of Patents?: Novelty Evaluation Based on the Correspondence between Patent Claim and Prior Art. | Claim-to-passage correspondence for kill-search; patent-specific labels do not establish scientific originality. Luna inspected the PDF and appendices. |
| M15 | Jeevan M. and Manjunatha B. N. (2026). AI-Based Patent Novelty Checker and Prior-Art Analysis Using Transformer Embeddings and FAISS. IJSREM, 10(07). | Retrieval/reporting prototype; dataset, accuracy and novelty-score calibration evidence insufficient for core-method adoption. Luna inspected the five-page PDF; article DOI not supplied. |
No upstream system was installed or benchmarked in this follow-up. Semantic matching and model critique provide evidence to inspect; they do not certify novelty. Deterministic checks can enforce records, schemas, versions and budgets, while semantic support and scientific value require substantive assessment.