Hive Trust · Benchmark Evidence

Review the reported results
and supporting evidence.

Review the reported metrics and what supports them. A signature authenticates a record when its signer is trusted. It does not establish that the benchmark was fair, reproducible or independently reviewed.

Hive ColonyIP protection shield Protected or Pending by Hive ColonyIP
Primitives in source checking…
Reported records checking…
Declared algorithm See record fields
Source timestamp age checking…

Source: checking. Signature and benchmark validity are not checked by these counters.

Reported Corpus Figures: Enterprise CASB Workload

SMSH v5 + smshPQMax: reported compression

These previously published corpus figures are retained as publisher reports, not accepted performance evidence. Their source is the June 3, 2026 website update, not an identified signed run artifact. Baseline revision, hardware, configuration, sample counts, uncertainty, retries and full compute costs are unknown. The stated recall bar is 99.5%.

Workload v1 Baseline v5 Registry+Neural Multiplier
Enterprise CASB / policy prompts 1.04x 9.15x 8.8x lift
Verbose / filler-heavy prompts 1.14x 5.09x 4.5x lift
RAG repetitive context 1.58x 3.41x 2.2x lift
Overall mean ~1.2x 5.49x 4.6x lift
Invariant Recall
99.78%
publisher-reported; run artifact unknown
Reported p50 latency
<25ms
unaccepted comparison: matched baseline, hardware and run artifact unknown
Signing
100%
publisher-reported receipt presence; executed ML-DSA verification not established

These corpus figures are not a head-to-head SOTA benchmark. The earlier 280x claim has been withdrawn because its matched run evidence was not located. The original website change records the figures above, but is not a benchmark run. The cards below identify individual records and their evidence gaps. smshPQMax product page →

Reported benchmarks

Cards identify whether records came from the service or a bundled fallback. Their status labels are publisher-reported, not an independent validation. The display does not verify signatures or reproduce a run.

Inference primitives: reported benchmark records

Other primitives: trust infrastructure

Voice primitives: STT/TTS compression and tamper detection

Badge designs

These are display examples, not certifications or awards to the listed primitives. The previously published statistical bar was n ≥ 500, |d| ≥ 0.3, p < 0.01. A badge or threshold does not replace signature checks, methodology review and a reproducible evidence package.

Hive Verified · badge design examples, not awards

Hive Platinum · badge design examples, not awards

How we benchmark

Requirements for an accepted comparison. Describing this method does not establish that every displayed record meets it.

Step 01

Pick the published SOTA

Select a relevant published baseline and freeze its version and configuration. The named comparison references include LLMLingua-2 for compression, NIST FIPS-204 for signatures, Llama-Guard for safety, self-consistency CoT for reasoning, DSPy for prompt compilation, Constitutional AI for factuality.

Step 02

Ensemble construction

The described Hive v2 ensemble includes a baseline candidate plus 3 to 4 Hive-specific strategies. Including a baseline does not guarantee the selector will choose the best output. Selection error, retries and extra compute can make an ensemble worse on quality, latency or total cost. Compare it on held-out inputs under the same budget and report losses as well as wins.

Step 03

Pre-registered evaluation

Freeze the dataset, sample size, metric, budget and decision criteria before a run. The linked pre-registration repository listed results as pending when checked on September 5, 2026. Pre-registration alone is not a completed reproducibility package: github.com/srotzin/xcalibur-evaluation.

Step 04

Cryptographic receipts

Verify each receipt against its trusted signer and exact signed bytes. The published query path is /v1/trust/benchmarks on receipts.thehiveryiq.com. A signature is not a check on experimental design. Acceptance also needs the run commit, dataset version, configurations, raw outputs, failures, full inference and selection costs, latency distribution, uncertainty and independent reproduction where claimed.

Result status

These are the publisher's record labels. A label or statistical threshold alone does not establish reproducibility, sound methodology or permission to make a public performance claim.

Publishable
Reported threshold: n ≥ 500, Cohen's d ≥ 0.3, p < 0.01. The full evidence package still needs review.
Preliminary
Publisher reports a minimum sample but insufficient effect size or significance. Do not treat this label as an accepted result.
Match
Publisher reports comparable correctness within its stated latency budget. Inspect the underlying run before accepting the comparison.

Inspect a benchmark record

The read-only API returned no benchmark records when checked on September 5, 2026. The cards therefore use the labeled bundled source when the live feed is empty or unavailable. Inspect that unchanged source.

curl -sS https://receipts.thehiveryiq.com/v1/trust/benchmarks/{record_id}

Verification requires the exact signed bytes and an approved issuer key from a separate trusted configuration. A record's pubkey_hex can support an internal consistency check, not establish issuer identity by itself. No approved benchmark issuer key is configured in this display. Some bundled records lack a signature field. Check supported formats before using the general receipt verifier; this link does not establish acceptance of the benchmark schema.

Open the verifier →
Hive ColonyIP

Hive primitives are Patent Pending. Provisional patents filed. The methodology, benchmarks, and receipts are open. The cryptographic primitives are protected.

Hive ColonyIP →