Your documents. Your hardware. Answers you can inspect.
Ask questions across agreements, reports, filings, and maintenance logs. KMS combines document retrieval with extracted relationships to assemble answers with source citations on customer-controlled hardware. Evaluate it on your own work in a 12-week pilot.
Exploring the model itself? View Valaris for Apple silicon (4-bit).

Bring a workflow and representative questions. Together we scope the corpus, hardware, access requirements, and success criteria before any documents are shared. The pilot is an evaluation proposal, not a promise of a particular score or speed.
Architecture illustration. Document workflows are the pilot focus; sensor inputs shown in the diagrams are development targets, not a claim of released sensor capabilities.
STRATUM extracts entities and relationships — subject, predicate, object — to build a graph alongside the document index. In the pilot, review whether those connections help answer your questions. Processing time depends on document size, format, and the selected hardware.
A covenant question touches a memo, a master agreement, and a position ledger. KMS resolves it by traversing the triple graph — seeded from vector hits, expanded breadth-first across extracted relationships, scored, then read by the local LLM with citations. Text search finds documents. Graph traversal finds answers.
KMS exposes retrieval over MCP (Model Context Protocol). Your agents query the same document and graph retrieval layer your analysts use. Identity, access boundaries, and audit requirements need to be assessed for each integration during the pilot.
Analyzing exposure: 42,000 bbl long option collar on Brent. MCTS search (10,000 paths) under Hormuz choke scenario yields tail VaR (P5) of -$4.2M over 10 trading days.
Risk-neutral density: Shift indicates +18% upside skew if Cape rerouting holds.
Recommended optimal hedge: Sell 15,000 bbl 30-day OTM calls ($84.50 strike) to finance 30-day OTM put spread ($72.00 / $66.00).
Private deployment and source citations do not, by themselves, establish access control. The pilot should document the intended users, permitted sources, and deployment boundary, then test access and audit behavior in that configuration. Authorization metadata or an LLM refusal is not proof that a permission boundary is enforced.
The diagram illustrates requirements to validate, not a security certification.
The existing reader comparison reports answer F1 on MuSiQue N=500, bare versus grounded with KMS. The Δ column shows the difference within that evaluation. These are research results on a fixed corpus, not expected scores for a customer deployment.
Source context: Sánchez & Dehnad, Selling the Meter, Proving the Outcome (MOSAIC, unpublished draft; Table D is cited in our research article). The figures below are retained as reported; the complete eight-reader run artifacts and energy-estimation method are not linked here for independent reproduction. This MuSiQue evaluation is separate from the 41-question explanation study in arXiv:2609.15319.
Answer F1 measures word overlap with reference answers on a 0–100 scale. A score of 49.2 is not 49.2% of questions answered correctly, and word overlap does not establish whether an explanation is supported. It is one measure to consider alongside source inspection and failure-case review. Recall@5 grades the search step: of the evidence passages actually needed to answer, the share that showed up in the system's top five results. First you have to find the evidence — that's Recall. Then you have to read it correctly — that's F1.
| # | LLM | Company | Origin | F1 Bare | F1 + KMS | Boost Δ | Deployment |
|---|---|---|---|---|---|---|---|
| 1 | Claude Fable 5 | Anthropic | 🇺🇸 | 49.1 | 65.0 | +15.9 | metered API |
| 2 | Claude Opus 4.8 | Anthropic | 🇺🇸 | 28.9 | 54.3 | +25.4 | metered API |
| 3 | ★ Toryx KMS + GemmaValaris | Toryx | 🇺🇸 | 15.6 | 49.2 | +33.6 | on-prem · air-gapped |
| 4 | Gemini 3.5 Flash | 🇺🇸 | 28.1 | 48.5 | +20.4 | metered API | |
| 5 | Llama 3.3 70B | Meta | 🇺🇸 | 19.0 | 46.4 | +27.4 | metered API |
| 6 | GPT-5.5 | OpenAI | 🇺🇸 | 16.1 | 46.2 | +30.1 | metered API |
| 7 | Kimi K3 | Moonshot AI | 🇨🇳 | 22.2 | 40.1 | +17.9 | metered API |
| 8 | GLM-5.2 | Z.ai | 🇨🇳 | 9.5 | 34.8 | +25.3 | metered API |
| LLM | F1 | 95% CI | Wh / answer (est.) |
|---|---|---|---|
| Opus 4.8 | 54.3 | [50.3, 58.3] | 0.229 |
| TORYX KMS (GemmaValaris) | 49.2 | [45.3, 53.1] | 0.00043–0.00058 |
| Llama 3.3 70B | 46.4 | [42.5, 50.4] | 0.010 |
| GLM-5.2 | 34.8 | [31.0, 38.7] | 0.067 |
Reported paired bootstrap, 10,000 resamples. Opus − GemmaValaris = +5.3 F1 [+2.2, +8.4] in this harness. CIs shown only where reported. Wh/answer values are estimates, not a measured customer power budget or total deployment cost.
Retrieval: Recall@5 = 82.7% [81.1, 84.3], MuSiQue N=1,000, nested CV. Research configuration on a non-commercial backbone (NV-Embed-v2); this is not a commercial-pipeline result. The full pipeline has not yet been rerun on a commercial backbone. In the separate bare-embedder comparison on the same 1,000 questions, Nemotron-3-Embed-8B scored 69.79 [68.13, 71.43] versus 69.55 for NV-Embed-v2, p=0.69. Source: arXiv:2608.16096.
The current pilot focuses on document ingestion, retrieval, extracted relationships, and cited answers. File formats and extraction quality should be checked on representative documents. Sensor adapters are a separate development track; thermal, acoustic, depth, and satellite retrieval are not included as released capabilities in this document pilot.
The sensor research direction explores adapters aligned to a multimodal backbone. Each modality needs its own data, evaluation, and release decision before it can be offered as a supported workflow.

Illustrative exhibit A — thermal retrieval research target, not a released capability

Illustrative exhibit B — passive acoustic retrieval, proof of concept planned

Illustrative exhibit C — satellite damage retrieval target, not a released capability
| Sensor Modality | Development Status | Primary Target Use Case |
|---|---|---|
| IR / Thermal Imaging | Research roadmap | Vehicle/platform thermal retrieval — defense, energy, maintenance |
| Underwater Audio | POC planned | Passive acoustic contact / whale-vessel framing — defense, maritime |
| Airborne Audio | POC planned | Drone / C-UAS acoustic retrieval — defense, law enforcement |
| Depth Maps | Research roadmap | Spatial reasoning and obstacle context — defense, maritime |
| Vibration / Acoustic Signatures | Planned | Equipment condition monitoring — maintenance, energy, aviation |
| Microscopy / Spectral | Planned | Biomedical image retrieval — pharma, clinical research |
| Satellite / Multispectral | Planned | Geospatial catastrophe & crop retrieval — finance, insurance |

Prototype shown: Mac Mini and NVIDIA DGX Spark systems. Proposed configurations can include Mac Studio; MacBook Pro options address mobile work. Final hardware depends on the pilot scope.
Discuss Apple hardware, NVIDIA DGX Spark, and MacBook Pro options against your corpus, models, mobility needs, and deployment boundary. The pilot should measure storage overhead, retrieval quality, and response time together. A photograph or raw storage figure is not a validated indexed-corpus capacity.
| Capability | Pilot review | Scope |
|---|---|---|
| Cross-document retrieval | Does the system find the evidence needed for your questions? | Document pilot |
| Relationship extraction | Are extracted entities and connections useful and accurate? | Document pilot |
| Cited answers | Do the cited passages support each answer and explanation? | Document pilot |
| Private deployment | Validate the selected hardware, network dependencies, and operating requirements. | Configuration-specific |
| Permissions and audit | Test the agreed access boundaries and audit behavior. | Integration-specific |
| Sensor question answering | Requires separate modality-specific evaluation before release. | Roadmap only |
Agree acceptance criteria before the evaluation. Research scores do not replace testing on your own corpus.
Bring the questions your documents should answer.
Start with a conversation about a 12-week pilot: one workflow, an agreed corpus, and evidence your reviewers can inspect. Read the research first, or explore the model as a separate technical resource.
Explore Valaris for Apple silicon on Hugging Face
info@toryx.ai