Repository navigation
Module llama cpp Future
Navigation: Home > Modules
Hinweis: Vage EintrΓ€ge ohne messbares Ziel, Interface-Spezifikation oder Teststrategie mit `` markieren.
Planned enhancements beyond v2.0.0. Core implementation in
src/llama_cpp/llama_cpp_plugin.cpp.
-
ILLMPlugininterface must remain the stable ABI; new capabilities added as optional methods with default implementations (e.g.,exportLoRA, streaming callback). - Stub mode (no model file,
loadModel("")) must remain functional after every enhancement. - Thread-safety guarantee (
std::mutex) must be maintained across all new code paths. - All calls to the real
LlamaWrappermust remain conditional onmodel_loaded_ == true. -
inference_count_anderror_count_must be incremented on every code path that calls or fails to call the underlying model.
| Interface | Consumer | Notes |
|---|---|---|
ILLMPlugin::generateStream(request, token_callback) |
Streaming API endpoint | New optional method; LlamaCppPlugin calls LlamaWrapper::generateStream()
|
ILLMPlugin::generateBatch(requests) |
Batch inference API | Returns vector<InferenceResponse>
|
ILLMPlugin::exportLoRA(lora_id) |
LoRA management API | Returns serialised adapter bytes |
ILLMPlugin::importLoRA(data, lora_id) |
LoRA management API | Deserialises and registers adapter |
Problem: generate() returns an echo stub in v2.0.0.
Solution: Add a real LlamaWrapper* member, initialised in loadModel(). Gate all
inference calls on THEMIS_ENABLE_LLAMA_CPP compile flag (consistent with
THEMIS_ENABLE_WHISPER pattern).
Inputs: InferenceRequest { prompt, max_tokens, temperature, top_p }.
Outputs: InferenceResponse { text, tokens_generated, latency_ms }.
Constraints: LlamaWrapper must be initialised before any generate() call;
double-init must be safe (unload + reload).
Errors: Model load failure β loadModel() returns false; generate on unloaded model
β error response.
Tests: Integration test with a tiny GGUF model in CI fixtures.
Perf target: β€ 200 ms for 50-token prompt on RTX 3090 equivalent.
Problem: generate() blocks until the entire response is generated.
Solution: Add generateStream(request, callback) to ILLMPlugin.
LlamaCppPlugin::generateStream() calls LlamaWrapper::generateStream() and forwards
each token to the callback on the calling thread.
Constraints: Callback must not throw; exceptions from the callback must be caught and
recorded as error_count_++.
Tests: 2 unit tests with a mock LlamaWrapper emitting 5 tokens.
Perf target: β€ 30 ms first-token latency on stub.
Problem: embed() returns a fixed 384-dim zero vector.
Solution: LlamaWrapper::embed() provides real embedding vectors when an embedding
model is loaded. LlamaCppPlugin::embed() will delegate to it.
Constraints: Embedding model may be different from the generation model; support loading both simultaneously. Tests: 3 unit tests: cosine similarity between related vs. unrelated texts. Perf target: β€ 5 ms per 512-token text on CPU.
Problem: exportLoRA returns empty and importLoRA returns false.
Solution: Serialize the LoRA weight matrices to a binary format (GGUF-compatible or
custom); importLoRA deserialises and hot-loads via LlamaWrapper::loadLoRA().
Security: Serialised LoRA bytes are validated (magic bytes, size bounds) before deserialisation to prevent injection attacks. Tests: Round-trip test: export β import β same weights.
Problem: supports_function_call is false in v2.0.0.
Solution: Add LlamaCppPlugin::callTool(request, tool_schema) using JSON schema
grammar-constrained generation (consistent with LLM module's grammar.cpp).
Constraints: Grammar validation required before compilation; recursive grammars bounded by depth limit (same constraint as LLM module). Tests: 5 unit tests with JSON schema fixtures.
Source: AI_ML_IMPACT_ASSESSMENT.md Β§7, Gap 1 (Severity: High/S1)
Status: β
Implemented (2026-04-21)
Problem: When LlamaCppPlugin::generate() is called without a loaded model
(wrapper_ == nullptr), it returned a stub response with success=true and text
"[stub:<prompt_prefix>]". Callers could not distinguish this from a real inference
result at the InferenceResponse level; silent incorrect output may propagate into
RAG pipelines and AQL results.
Implemented changes:
-
generate()now returnssuccess=false+error_message="Model not loaded β call loadModel() before generate()"whenwrapper_is nullptr andTHEMIS_LLAMA_CPP_STUB_MODEis not defined. - Test builds define
THEMIS_LLAMA_CPP_STUB_MODEvia CMakeLists to preserve the echo stub for existing tests (D2/D3/N6 groups). - New Group O tests (O1..O3) added to
src/llama_cpp/tests/test_llama_cpp_plugin.cppverify the production error path and stub-mode compatibility.
Inputs: InferenceRequest (unchanged); wrapper_ state (nullptr vs. loaded).
Outputs: InferenceResponse { success=false, error_message }.
Constraints: Existing unit tests that rely on the stub response use THEMIS_LLAMA_CPP_STUB_MODE.
Tests: Group O tests (O1..O3) in src/llama_cpp/tests/test_llama_cpp_plugin.cpp.
Perf target: No performance impact (error path only).
- All new inference paths must pass through
PolicyEngine::checkInferencePermission()before queueing (consistent with LLM module security policy). -
importLoRAmust validate size bounds before heap allocation. -
generateStreamcallbacks must never receive pointers to stack-allocated token data that may be invalidated after the streaming call returns.
ThemisDB 1.9.0-beta Β· Home Β· Module-Index Β· GitHub Β· Issues
ThemisDB 1.9.0-beta Β· Home Β· Wiki-Index Β· Module-Index Β· FAQ Β· Quick-Reference Β· GitHub Β· Issues Β· Discussions Β· License
- Home
- Hero Articles
- All Wiki Pages
- FAQ
- Edition Comparison
- Repository README
- Changelog
- Roadmap
- Versioning
- Integration Mapping
- Overview
- Readme
- Appendix D Feature Status
- Appendix E Incident Runbooks
- Appendix F AQL Cheatsheet
- Appendix G Configuration
- Appendix H Glossary
- Appendix I Troubleshooting
- Appendix Literatur
- Chapter 00 Genesis
- Chapter 01 Introduction
- Chapter 02 Architecture
- Chapter 03 Multimodel
- Chapter 04 Installation
- Chapter 05 Relational
- Chapter 06 Graph
- Chapter 07 Document
- Chapter 08 Storage Layer
- Chapter 08 Vector
- Chapter 09 Timeseries
- Chapter 10 Enterprise
- Chapter 11 Realtime
- Chapter 12 Computervision
- Chapter 13 Fulltext
- Chapter 14 Geospatial
- Chapter 15 Analytics
- Chapter 16 Ml
- Chapter 16 Sharding
- Chapter 17 LLM Integration
- Chapter 17 Scaling
- Chapter 18 HA
- Chapter 18 Ml
- Chapter 19 Monitoring
- Chapter 19 Monitoring Observability
- Chapter 20 Backup
- Chapter 20 Performance
- Chapter 21 Auth
- Chapter 21 Performance
- Chapter 22 Clients
- Chapter 22 Encryption
- Chapter 23 Testing Qa
- Chapter 24 Ai Ethics
- Chapter 25 Devops Infrastructure
- Chapter 26 Migration Legacy
- Chapter 27 Troubleshooting
- Chapter 28 AQL Reference
- Chapter 29 Analytics Process Mining
- Chapter 30 Deployment Operations
- Chapter 31 API Protocols
- Chapter 32 API Design Rest Principles
- Chapter 32 AQL Oop Implementation
- Chapter 33 Best Practices
- Chapter 34 Query Optimization
- Chapter 35 Data Modeling Patterns
- Chapter 36 Security Hardening
- Chapter 37 Ecosystem Integration
- Chapter 38 Observability Sre
- Chapter 39 Performance Tuning Cookbook
- Chapter 40 Data Governance Compliance
- Chapter 41 Hands On Labs
- Chapter 42 Docs Assistant Usage
- Chapter MVCC Hlc
- Cover
- Cover Book
- Index
- Preface
- Test Links Example
- Batch Operations
- Best Practices
- CRUD Tutorial
- Custom Document Ingestion
- Getting Started Tutorial
- Interactive Examples
- Schema Design
- Video Tutorials
- AQL Reference
- AQL Examples
- AQL Overview
- AQL Feature Roadmap
- AQL Geospatial Guide
- AQL LLM Migration Guide
- AQL API
- AQL Grammar (EBNF)
- AQL Root Overview
- AQL Examples (root)
- API Reference
- API Module README
- OpenAPI Overview
- Client SDK Overview
- SDK Overview
- Operations
- Operations Overview
- Operations Runbook
- Operations Handbook
- ThemisCtl Admin Guide
- Pipeline E2E SOPs
- Docker Overview
- Docker Hub README
- Helm Overview
- Packaging Overview
- Operator Overview
- Security Policy
- Production Hardening Checklist
- Security Hardening Guide
- Encryption Key Management
- Access Control Framework
- Zero Trust Policy
- API Authentication & Authorization
- HSM Production Setup
- PKCS11 Integration
- DSGVO / SOC2 Checklist
- Access Model Runbooks
- Access Model Dashboard
- Maturity Automation Runbook
- Access Review Automation
- Access Model Dashboard
- Access Model Runbooks
- Rights Revocation
- Dr Checklists
- Dr Testing
- Incident Response Playbook
- Incident Response Testing
- GPU Oom Recovery
- Grammar Debugging
- Metrics Scrape Troubleshooting
- Model Swap Procedure
- Quota Tuning
- Subagent Deployment
- Logging Configuration
- Content Model
- Crypto & Keys
- Feature Flags Reference
- Modular Architecture Roadmap
- Modularization Guide
- Module Architecture Index
- PostgreSQL Wire Protocol
- Query Scheduling
- Raft Consensus Design
- Resource Pooling
- Source Directory Guide
- Unified Access Model
- E1 001 Layered Retrieval Design
- E1 002 Ann Abstraction Strategy
- E1 003 Tensor Summary Types
- E1 004 Lora Package Distinction
- E1 005 Model Switch Compatibility
- E1 006 Federated Tensor Summaries
- E2 001 Evaluation Framework Design
- E2 002 Hardware Profile Strategy
- E2 003 Query Planner Routing Model
- E2 004 Approximation Governance Rules
- E2 005 Cross Layer Fallback Confidence Policy
- E3 001 Distributed Tensor Design
- E3 002 Manifest Coordination Strategy
- E3 003 Recovery And Erasure Choice
- E3 004 Tensor Fabric Infrastructure
- Contributing
- Contributing (root)
- Code of Conduct
- Support
- Maintainers
- CTest Guide
- Build Quick Reference
- Developer Wiki Index
- Build / Test / CI
- Module Index
- Branching Strategy
- Release Strategy
- CI Policy Gates Wave C
- Disabled Stub Policy
- Docs PR Policy
- GA Promotion Sign Off
- Github Milestones Setup
- Governance Policies Phase1
- GPU Self Hosted Runner Requirements
- Hardening Phase 1 2 Summary 2026 09 23
- Maturity Claim Verification Checklist
- Maturity Evidence Registry
- Merge Gate Bot Config
- Merge Gate Status Live
- Phase 1 Closure Report
- Phase 1 Infrastructure Deployment
- Phase 1 Infrastructure Deployment Complete
- Phase 3 Baseline Capture
- Phase 3 Refinement Spec
- Phase 4 Sign Off And Closure
- Phase Closure Policy
- Phase Dependency Graph
- Phase3 Enforcement Runbook
- Plugin Submodule Rollback
- PR Version Targeting
- PR Version Targeting Backfill
- Production Ready 2026 Delivery Plan
- Publish Workflow Audit 2026 09 23
- Query Module Status
- Readme
- Release Governance
- Release Promotion Gate Policy
- Release Validation Checklist
- Root Hygiene Policy
- SBOM Approved Versions
- Security Compliance Audit Report 2026 08 10
- Security Module 5671 Evidence Summary
- Sharding P6 Residual Risk Acceptance
- Sourcecode Compliance Governance
- Src Module Documentation Compliance 2026 09 20
- Updates Development Status Sign Off
- Wave C Implementation Complete
- Wave C Implementation Plan
- Wave C Ml Exit Gate Sign Off
- Wave C Policy Gate Evidence
- Wiki Publish Tracking Guide
- Blob Storage
- Cuda
- Ethics Ai
- Exporters
- Huggingface
- Image Analysis
- Importers
- RPC
- Scraper
- Themisdb Ai Watermark Detector
- User Storage Encrypted
- Chimera Architecture
- Chimera Future
- Chimera Readme
- Chimera Roadmap
- Covina Fastapi Ingestion Architecture
- Covina Fastapi Ingestion Future
- Covina Fastapi Ingestion Roadmap
- Vcc Base Architecture
- Vcc Base Future
- Vcc Base Roadmap
- Vcc Clara Ingestion Architecture
- Vcc Clara Ingestion Future
- Vcc Clara Ingestion Roadmap
- Vcc Veritas Architecture
- Vcc Veritas Future
- Vcc Veritas Roadmap
- 01 Hello World
- 02 Todo App
- 03 Contact Manager
- 04 Inventory System
- 05 Time Series Monitor
- 06 Graph Social Network
- 07 Vector Search Documents
- 08 Dms Erp System
- 09 Iot Sensor Network
- 10 Drone Image Analysis
- 11 Blog Wiki
- 12 Expense Tracker
- 13 Recipe Manager
- 14 Ecommerce Catalog
- 15 Event Management
- 16 Kanban Board
- 17 Crm
- 18 Realtime Chat
- 19 Recommendation Engine
- 20 Smart Home
- 21 Coding Platform
- 22 AQL Diagram Tool
- 23 Traveling Salesman
- 24 Moral Philosophy Debates
- API Versioning
- Distributed Sharding
- Feedback Plugins
- Geo
- Gnn
- Image Analysis
- Legal Lora Training
- LLM
- Lora Sync
- Migration
- Nlp
- Performance
- Railway
- Replication
- Rope Visualization
- Sample Product Config
- Security
- Client SDK Overview
- Quickstart
- Sdk Enhancements
- Sdk Implementation Summary
- Test Suite Readme
- Go
- Java
- Javascript
- Php
- Python
- Ruby
- Rust
- Typescript
- 01 Grundlegende Operationen
- 02 AQL Queries
- 03 Graph Daten
- 04 Multimodell Anwendung
- 01 Quickstart Guide
- 02 AQL Referenz Kurzuebersicht
- 03 Datenmodellierung Guide
- 04 Uebungsaufgaben
- 05 Best Practices Guide
- Training Documents
- Training Overview
- 01 Einfuehrung Und Uebersicht
- 02 Datenmodelle Und Architektur
- 03 AQL Abfragesprache
- 04 Installation Und Setup
- 05 Anwendungsbeispiele
- Training Presentations
- Dependencies Readme
- Processmonitor Readme
- Themis.admintools.shared Readme
- Themis.aqlquerybuilder Readme
- Themis.aqlquerybuilder Roadmap
- Themis.auditlogviewer Readme
- Themis.auditlogviewer Roadmap
- Themis.classificationdashboard Readme
- Themis.classificationdashboard Roadmap
- Themis.compliancereports Readme
- Themis.compliancereports Roadmap
- Themis.gisviewer.controlpanel Readme
- Themis.gisviewer.controlpanel Roadmap
- Themis.impactanalysisviewer Readme
- Themis.impactanalysisviewer Roadmap
- Themis.ingestiontool Readme
- Themis.ingestiontool Roadmap
- Themis.keyrotationdashboard Readme
- Themis.keyrotationdashboard Roadmap
- Themis.piimanager Readme
- Themis.piimanager Roadmap
- Themis.retentionmanager Readme
- Themis.retentionmanager Roadmap
- Themis.sagaverifier Readme
- Themis.sagaverifier Roadmap
- Themis.usbadmintool Readme
- Themis.usbadmintool Roadmap
- Architecture Generator Readme
- CI Readme
- CI Roadmap
- Compiler Diagnostics Readme
- Compiler Diagnostics Roadmap
- Completion Readme
- Copilot Ollama Router Readme
- Copilot Ollama Router Roadmap
- Gnn Readme
- Gnn Roadmap
- Rope Visualizer Readme
- Rope Visualizer Roadmap
- Tco Calculator Readme
- Tco Calculator Roadmap
- Tests Readme
- Tests Roadmap
- Themis Config Wx Readme
- Themis Docs Builder Readme
- Wikipedia Ingestion Readme
- Ai Metadata And Provenance
- Build / Test / CI
- Governance And Roadmap
- Developer Wiki Index
- Module Direct Doxygen Check
- Module Doxygen Baseline Summary
- Module Doxygen Batch
- Module Doxygen Coverage Summary
- Module Doxygen Smoke Summary
- Modules And Apis
- Retrieval Direct Doxygen Check
- Soll Ist Gap Summary
- Wiki Delta Report