KMS · KNOWLEDGE MANAGEMENT SYSTEM

Private AI that connects your documents—and shows its evidence.

Your documents. Your hardware. Answers you can inspect.

Ask questions across agreements, reports, filings, and maintenance logs. KMS combines document retrieval with extracted relationships to assemble answers with source citations on customer-controlled hardware. Evaluate it on your own work in a 12-week pilot.

Exploring the model itself? View Valaris for Apple silicon (4-bit).

Toryx prototype: two NVIDIA DGX Sparks, a Mac Mini, networking and an NVMe enclosure
Apple + NVIDIA DGX Sparkcustomer-controlled
Mac Mini prototype shown. Proposed Mac Studio configuration; MacBook Pro options for mobile work.
START WITH YOUR DOCUMENTS

A 12-week pilot, with a question worth answering.

Bring a workflow and representative questions. Together we scope the corpus, hardware, access requirements, and success criteria before any documents are shared. The pilot is an evaluation proposal, not a promise of a particular score or speed.

  1. 1. Define the evaluationAgree which questions, evidence, and review criteria matter to your team.
  2. 2. Test the document workflowEvaluate retrieval, cited answers, and failure cases on the agreed corpus and configuration.
  3. 3. Review the resultUse the evidence and operational findings to decide whether to expand.

INGESTION PIPELINE & ARCHITECTURE

From raw file to cited answer — inside your perimeter.

YOUR PERIMETERFiles & sensorsdocuments & feedsSTRATUMstructure + embedSovereign storegraph + vectorsScoringevidence scoredLocal LLMGemmaValaris-26BCited answerzero marginal
No stage calls a cloud API.

Architecture illustration. Document workflows are the pilot focus; sensor inputs shown in the diagrams are development targets, not a claim of released sensor capabilities.

  • Extract entities and relationships from documents to support cross-document questions.
  • Combine document retrieval with graph relationships, then inspect the supporting sources.
  • Evaluate ingestion quality, index updates, and retrieval coverage on the pilot corpus.
SYSTEM HUB & SPOKE SCHEMATICCollectorsfeeds & sensorsSTRATUM enginetriple extractionSovereign storegraph + vectorsGraph traversalbreadth-firstScoringevidence scoredGemmaValaris LLMon-prem fine-tuneMCP interfaceModel Context ProtocolAgent fleetgoverned execution

STRATUM ENGINE

STRATUM — automatic ontology at the embedding layer.

STRATUM extracts entities and relationships — subject, predicate, object — to build a graph alongside the document index. In the pilot, review whether those connections help answer your questions. Processing time depends on document size, format, and the selected hardware.

Processing time:Measure on your corpus
Integration scope:Defined in the pilot
IP portfolio:20 patent filings
STRATUM extraction exampleILLUSTRATIVE
$ stratum ingest --file="Credit_Agreement_2023.pdf" --extract-triples
Example relationships for review
<Meridian_Logistics> → <hasLeverageRatio> → <3.4x>
<Master_Agreement> → <hasCovenantCeiling> → <3.5x>
<Covenant_Breach> → <triggersCollateralStepUp> → <10_Business_Days>
Illustrative data — fictional entities.

Rotating 3D point cloud where modalities mix by topic, showing a three-hop geopolitical query touching satellite AIS data, an internal memo, and a news item

One retrieval space · points cluster by meaning, not modality · drag to rotate
query: "does the Hormuz escalation affect the fundamentals of our long positions with supply chains outside MENA?"
Illustrative data — fictional positions.
GRAPH TRAVERSAL ENGINE

Questions that hop.

A covenant question touches a memo, a master agreement, and a position ledger. KMS resolves it by traversing the triple graph — seeded from vector hits, expanded breadth-first across extracted relationships, scored, then read by the local LLM with citations. Text search finds documents. Graph traversal finds answers.


AGENTIC INTERFACE · MCP

Built for agents, not just analysts.

KMS exposes retrieval over MCP (Model Context Protocol). Your agents query the same document and graph retrieval layer your analysts use. Identity, access boundaries, and audit requirements need to be assessed for each integration during the pilot.

2 · Act — agentic quant loop
@toryx-valaris — MCTS and VaR on the exposed position, and the optimal hedge.

Analyzing exposure: 42,000 bbl long option collar on Brent. MCTS search (10,000 paths) under Hormuz choke scenario yields tail VaR (P5) of -$4.2M over 10 trading days.

Risk-neutral density: Shift indicates +18% upside skew if Cape rerouting holds.

Recommended optimal hedge: Sell 15,000 bbl 30-day OTM calls ($84.50 strike) to finance 30-day OTM put spread ($72.00 / $66.00).

Illustrative data — fictional entities and prices.

usercleared · level Arecord 041LEVEL Arecord 187LEVEL Brecord 312LEVEL Xillustrative boundaries · verify in pilot
GOVERNANCE · PILOT REQUIREMENTS

Define who should see what. Then test it.

Private deployment and source citations do not, by themselves, establish access control. The pilot should document the intended users, permitted sources, and deployment boundary, then test access and audit behavior in that configuration. Authorization metadata or an LLM refusal is not proof that a permission boundary is enforced.

The diagram illustrates requirements to validate, not a security certification.


EVALUATION · MUSIQUE N=500 BAKE-OFF

What the document benchmarks measured.

The existing reader comparison reports answer F1 on MuSiQue N=500, bare versus grounded with KMS. The Δ column shows the difference within that evaluation. These are research results on a fixed corpus, not expected scores for a customer deployment.

Source context: Sánchez & Dehnad, Selling the Meter, Proving the Outcome (MOSAIC, unpublished draft; Table D is cited in our research article). The figures below are retained as reported; the complete eight-reader run artifacts and energy-estimation method are not linked here for independent reproduction. This MuSiQue evaluation is separate from the 41-question explanation study in arXiv:2609.15319.

What these numbers mean — in plain English

Answer F1 measures word overlap with reference answers on a 0–100 scale. A score of 49.2 is not 49.2% of questions answered correctly, and word overlap does not establish whether an explanation is supported. It is one measure to consider alongside source inspection and failure-case review. Recall@5 grades the search step: of the evidence passages actually needed to answer, the share that showed up in the system's top five results. First you have to find the evidence — that's Recall. Then you have to read it correctly — that's F1.

#LLMCompanyOriginF1 BareF1 + KMSBoost ΔDeployment
1Claude Fable 5Anthropic🇺🇸49.165.0+15.9metered API
2Claude Opus 4.8Anthropic🇺🇸28.954.3+25.4metered API
3★ Toryx KMS + GemmaValarisToryx🇺🇸15.649.2+33.6on-prem · air-gapped
4Gemini 3.5 FlashGoogle🇺🇸28.148.5+20.4metered API
5Llama 3.3 70BMeta🇺🇸19.046.4+27.4metered API
6GPT-5.5OpenAI🇺🇸16.146.2+30.1metered API
7Kimi K3Moonshot AI🇨🇳22.240.1+17.9metered API
8GLM-5.2Z.ai🇨🇳9.534.8+25.3metered API

Reported confidence intervals (MOSAIC research configuration)

LLMF195% CIWh / answer (est.)
Opus 4.854.3[50.3, 58.3]0.229
TORYX KMS (GemmaValaris)49.2[45.3, 53.1]0.00043–0.00058
Llama 3.3 70B46.4[42.5, 50.4]0.010
GLM-5.234.8[31.0, 38.7]0.067

Reported paired bootstrap, 10,000 resamples. Opus − GemmaValaris = +5.3 F1 [+2.2, +8.4] in this harness. CIs shown only where reported. Wh/answer values are estimates, not a measured customer power budget or total deployment cost.

Retrieval: Recall@5 = 82.7% [81.1, 84.3], MuSiQue N=1,000, nested CV. Research configuration on a non-commercial backbone (NV-Embed-v2); this is not a commercial-pipeline result. The full pipeline has not yet been rerun on a commercial backbone. In the separate bare-embedder comparison on the same 1,000 questions, Nemotron-3-Embed-8B scored 69.79 [68.13, 71.43] versus 69.55 for NV-Embed-v2, p=0.69. Source: arXiv:2608.16096.


DOCUMENT CAPABILITIES · SENSOR ROADMAP

Documents first. Sensor research next.

The current pilot focuses on document ingestion, retrieval, extracted relationships, and cited answers. File formats and extraction quality should be checked on representative documents. Sensor adapters are a separate development track; thermal, acoustic, depth, and satellite retrieval are not included as released capabilities in this document pilot.

Input → Evaluation scope
Text & documentsPilot focus: search across documents and inspect cited answers
Scanned PDFs & slidesEvaluate document extraction and retrieval on representative files
Code & schemasConfirm supported formats and retrieval needs during scoping
Thermal, audio, depth & satellite inputsRoadmap: separate adapter research and validation, not released pilot features

The sensor research direction explores adapters aligned to a multimodal backbone. Each modality needs its own data, evaluation, and release decision before it can be offered as a supported workflow.


SENSOR ADAPTER ROADMAP & EXHIBITS

Multimodal sensor adapter roadmap.

Thermal IR Frame

Illustrative exhibit A — thermal retrieval research target, not a released capability

Passive Acoustic Array

Illustrative exhibit B — passive acoustic retrieval, proof of concept planned

Satellite Damage Grid

Illustrative exhibit C — satellite damage retrieval target, not a released capability

Sensor ModalityDevelopment StatusPrimary Target Use Case
IR / Thermal ImagingResearch roadmapVehicle/platform thermal retrieval — defense, energy, maintenance
Underwater AudioPOC plannedPassive acoustic contact / whale-vessel framing — defense, maritime
Airborne AudioPOC plannedDrone / C-UAS acoustic retrieval — defense, law enforcement
Depth MapsResearch roadmapSpatial reasoning and obstacle context — defense, maritime
Vibration / Acoustic SignaturesPlannedEquipment condition monitoring — maintenance, energy, aviation
Microscopy / SpectralPlannedBiomedical image retrieval — pharma, clinical research
Satellite / MultispectralPlannedGeospatial catastrophe & crop retrieval — finance, insurance

Toryx prototype with Apple Mac Mini, NVIDIA DGX Spark systems, networking, and storage

Prototype shown: Mac Mini and NVIDIA DGX Spark systems. Proposed configurations can include Mac Studio; MacBook Pro options address mobile work. Final hardware depends on the pilot scope.

HARDWARE · SIZED TO THE WORKFLOW

Choose the configuration after defining the work.

Discuss Apple hardware, NVIDIA DGX Spark, and MacBook Pro options against your corpus, models, mobility needs, and deployment boundary. The pilot should measure storage overhead, retrieval quality, and response time together. A photograph or raw storage figure is not a validated indexed-corpus capacity.


PILOT CAPABILITY CHECKLIST

What to evaluate before expanding.

CapabilityPilot reviewScope
Cross-document retrievalDoes the system find the evidence needed for your questions?Document pilot
Relationship extractionAre extracted entities and connections useful and accurate?Document pilot
Cited answersDo the cited passages support each answer and explanation?Document pilot
Private deploymentValidate the selected hardware, network dependencies, and operating requirements.Configuration-specific
Permissions and auditTest the agreed access boundaries and audit behavior.Integration-specific
Sensor question answeringRequires separate modality-specific evaluation before release.Roadmap only

Agree acceptance criteria before the evaluation. Research scores do not replace testing on your own corpus.

Bring the questions your documents should answer.

Start with a conversation about a 12-week pilot: one workflow, an agreed corpus, and evidence your reviewers can inspect. Read the research first, or explore the model as a separate technical resource.

Explore Valaris for Apple silicon on Hugging Face

info@toryx.ai