Corrected replay + historical run · 2026-09-12
One hundred searches.
Every miss counted.
The original responses stay frozen. The current Fast-to-Standard consistency rule is replayed over them, so the corrected metrics are traceable without inventing a new benchmark run.
Corrected fixed-mode replay
The frozen Fast prefix is preserved, unique frozen Standard results are appended to ten, and the original scoring formula is applied unchanged. No provider was called again.
| Orbita mode | Quality | Hit@1 | Hit@5 | Hit@10 | MRR | nDCG@10 | Status |
|---|---|---|---|---|---|---|---|
| Standard | 80.01 | 65% | 88% | 91% | 0.7363 | 0.7789 | deterministic replay |
| Fast | 79.37 | 65% | 88% | 88% | 0.7317 | 0.7686 | unchanged frozen result |
Boundary. These are corrected rank metrics from stored responses, not a new live run. Latency was not replayed and remains historical until the next frozen benchmark.
Historical leaderboard
The original values are retained rather than rewritten after the product correction. Composite rank is diagnostic here—not a promise that Fast is the stronger mode.
| Provider / mode | Quality | Hit@1 | Hit@5 | Hit@10 | Hit@10 95% CI | MRR | nDCG@10 | p50 | p95 | Errors |
|---|---|---|---|---|---|---|---|---|---|---|
| Orbita Fast | 79.37 | 65% | 88% | 88% | 80.19–93.00% | 0.7317 | 0.7686 | 836 ms* | 5,672 ms* | 0 |
| Orbita Standard | 76.06 | 61% | 81% | 88% | 80.19–93.00% | 0.6990 | 0.7427 | 5,394 ms* | 16,746 ms* | 0 |
| Exa Fast | 68.23 | 62% | 67% | 67% | 57.31–75.44% | 0.6400 | 0.6476 | 624 ms | 970 ms | 0 |
| Linkup Standard | 52.92 | 38% | 52% | 61% | 51.20–69.98% | 0.4498 | 0.4877 | 1,327 ms | 2,312 ms | 0 |
| Tavily Basic | 35.97 | 18% | 37% | 40% | 30.94–49.80% | 0.2574 | 0.2923 | 2,065 ms | 4,402 ms | 0 |
Mode consistency. This run exposed a regression where a broader fixed mode could reorder a strong Fast result downward. That behavior has been corrected; fresh rank and latency numbers require a new frozen run. External and Orbita timings here are not directly comparable.
By query stratum
The factual subset is harder and smaller. The navigation subset asks for an official page by title, so it should not be read as open-ended research quality.
| Provider / mode | Factual quality | Factual Hit@10 | Navigation quality | Navigation Hit@10 |
|---|---|---|---|---|
| Orbita Fast | 75.82 | 77.27% | 80.37 | 91.03% |
| Orbita Standard | 75.00 | 81.82% | 76.36 | 89.74% |
| Exa Fast | 59.45 | 59.09% | 70.70 | 69.23% |
| Linkup Standard | 24.41 | 36.36% | 60.97 | 67.95% |
| Tavily Basic | 33.55 | 36.36% | 36.65 | 41.03% |
Method
Every target URL was fixed and confirmed before execution. The corpus size, storage layout and internal index statistics are intentionally not public.
Quality formula
Hit@1 25% + Hit@5 20% + Hit@10 15% + MRR 20% + nDCG@10 10% + successful delivery 10%.
Exact modes
Orbita Fast 20→5; Orbita Standard 100→10; Exa Fast; Tavily Basic; Linkup Standard. No failed request was retried.
What this does not prove
These limits stay beside the result so a benchmark cannot quietly become a larger marketing claim.
- The factual stratum is the active subset of the prior pre-authored benchmark.
- The navigation stratum measures retrieval of a known official page by title and is easier than open-ended research.
- Orbita deliberately contained every target; external providers searched their own indexes. This is not a web-coverage comparison.
- Raw responses and misses remain frozen. A miss or failed delivery stays in the denominator.