System report · completed 2026-09-13

Corpus intake.
What we can prove.

A public, outcome-focused account of corpus intake, retrieval quality, account experience and remaining risks. Database size and proprietary implementation details are intentionally omitted.

unique intake set87.2% searchableexplicit exclusionsreviewed sources221-test snapshot

Historical snapshot. The figures below describe the completed 2026-09-12/13 run. Reliability and pricing changed later, but the original measurements remain unchanged.

87.2%of the frozen intake set became searchable
12.8%ended as explicit, inspectable exclusions
100%of accepted input URLs were unique at intake
$0.01769confirmed processing spend for the batch

Corpus expansion

The batch used reviewed public sources, normalized duplicate URLs and respected publisher access rules. Internal ingestion stages and infrastructure are not part of the public report.

MeasureResultPublic meaning
Input uniqueness100%Every accepted URL in the frozen intake set was unique.
Searchable outcome87.2%Pages completed intake and became available to retrieval.
Explicit exclusion outcome12.8%No failed page was silently counted as searchable.
Unique content after deduplication98.2%Nearly all searchable intake results represented distinct content.
Publication checks100%Every searchable intake record passed the required publication checks.
Confirmed processing spend$0.017690044Recorded processing cost for this frozen run.

Every exclusion accounted for

The public view reports the distribution of outcomes without revealing corpus size, internal thresholds or recovery procedures.

Could not read safely90.9%
Unavailable upstream4.9%
Publisher/access policy4.2%

Subsequent status. Reliability was improved after this frozen run. Unavailable pages and publisher restrictions still remain explicit exclusions.

Search100

The corrected rank view replays the current Fast-to-Standard consistency rule over unchanged stored responses. The original frozen-run table remains below for auditability.

Corrected fixed-mode replay

Orbita modeQualityHit@1Hit@5Hit@10MRRnDCG@10
Standard80.0165%88%91%0.73630.7789
Fast79.3765%88%88%0.73170.7686

Replay, not fabrication. No provider was called again and no score was manually increased. The original scoring formula was applied after the corrected mode contract combined the frozen outputs.

Original frozen run

Provider / modeQualityHit@1Hit@5Hit@10MRRnDCG@10p50p95
Orbita Fast79.3765%88%88%0.73170.7686836 ms*5,672 ms*
Orbita Standard76.0661%81%88%0.69900.74275,394 ms*16,746 ms*
Exa Fast68.2362%67%67%0.64000.6476624 ms970 ms
Linkup Standard52.9238%52%61%0.44980.48771,327 ms2,312 ms
Tavily Basic35.9718%37%40%0.25740.29232,065 ms4,402 ms

Read narrowly. A fresh live run is still required to confirm post-correction latency and repeatability; the original table is not silently relabelled. Read the full benchmark methodology →

Factual subset

Standard found 18 of 22 known factual sources in its first ten; Fast found 17 of 22. That is the meaningful depth signal retained from this run.

Execution notes

All calls were first attempts. One AP-title Standard outlier took about four minutes and remains in the frozen record. External published-rate estimate: Exa $0.70, Tavily $0.80, Linkup $0.50.

Accounts, billing and API

The staging flow was exercised end to end without pretending that payment infrastructure already existed.

Verification at completion

The original report froze the state that existed when the run ended; later repository tests are tracked separately.

Automated

221/221 repository tests passed sequentially, followed by 3/3 final changed-surface tests. Internal service topology is intentionally omitted.

Visual

Search100 chart, privacy block, Overview, API keys and Billing/payment placeholder passed the final visual check.

Remaining blockers

The report ended with these risks visible rather than burying them beneath the benchmark score.