Ghost Search — MCP-native, Tor-routed search layer for RAG (seeking research collaborators)
## **Ghost Search MCP — Production Hardening Complete**
This is an update to the previous post. Every item from the “What’s next” section has been shipped, plus additional hardening work. Here’s where each landed.
### **Replay delta — root-caused and fixed**
The -0.1191 nDCG@10 gap between the benchmark evaluator (`eval-ranking.ts`) and the production ranker (`search.ts`) had two root causes:
1. **Quality filter replaced RRF scores entirely.** `applyQualityFilter()` sorted by `qualityScore`, discarding the RRF ordering. This accounted for ~-0.06 nDCG.
2. **Per-engine info-quality sort was missing in production.** `eval-ranking.ts` sorted per-engine results by snippet informativeness before assigning RRF ranks; `rankProduction()` did not. This accounted for ~-0.06 nDCG.
Fix: the quality filter is now a modifier on RRF scores, not a replacement. Penalized results get `rrfScore * 0.1`; non-penalized results get `rrfScore + qfScore * 0.01`. The per-engine info-quality sort was added to both `rankProduction()` and the live `executeSearch()` path.
Metric | Before fix | After fix
---|---|---
Eval nDCG@10 | 0.4080 | 0.3900
Production nDCG@10 | 0.2889 | 0.3944
Delta | -0.1191 | +0.0044
The drift is eliminated. The remaining +0.0044 is the quality filter penalty working as intended — production is marginally stricter than the evaluator on duplicate/mirror results, which is the correct direction.
The drift analysis tool (
analyze-ranking-drift.ts) remains in therepo. Anyfuturerankingchange can be replayed against the frozenv1corpus and measuredimmediately.
### **Engine query contracts — all 13 engines covered**
The previous post noted 3 of 13 engines had documented contracts (Ahmia, Torch, TorDex). All 13 now have `QueryContract` definitions with `maxQueryLength`, `maxTerms`, `matchMode`, `booleanOperators`, and `phraseQuotes` constraints. The query adapter (
query-adapter.ts) normalizes queriesperengine beforedispatch—truncating length, limitingterms, stripping Booleanoperators andphrase quotes when unsupported.
Contract values are provisional empirical defaults from live Tor probing, not authoritative engine specifications. HTTP 400 responses are retried once with 2s backoff (transient 400s observed on multiple engines). TorDex’s contract was revised twice after live probing showed error_400 is transient rather than length-based — 10 terms / 84 chars confirmed working on retry.
Contract coverage tests (
engine-contracts.test.ts) verify everyengine has a contractand that`adaptQuery`respectseach constraint.
### **Query planner — implemented**
The federated acquisition framing from the previous post is now shipped. When `queryPlanner: true` is passed to `ghost_search` or `ghost_batch_search`:
1. An LLM generates 2-3 short atomic sub-queries from the original query
2. Each sub-query is adapted per engine contract via the existing query adapter
3. Sub-queries are sent to engines as additional search waves via `runEngineWave()`
4. Results are fused into RRF with trust weight = 1.0 (low, experimental)
5. Single response — no protocol break, no streaming
The LLM cascade is OpenRouter DeepSeek V4 Flash → NVIDIA NIM → Ollama → raw query (fail-soft). If every LLM is unavailable, the search proceeds with the original query. The planner is off by default in the production-v1 profile.
Latency impact: adds one LLM call (~1-3s OpenRouter) plus 2-3 additional engine waves (~20-60s over Tor). This is opt-in and not recommended for interactive use without the operator accepting the latency trade-off.
Replay against frozen v1 with planner OFF: nDCG@10 = 0.3944 (no regression). Live Tor test with planner ON remains as a manual operator-driven task — the replay framework cannot measure it because the planner generates different sub-queries each run.
### **Additional shipped work**
**Structured audit logging (PR #25).** Every `ghost_search` call produces one structured JSON line on stderr: tool name, query, engine count, result count, latency, engine health. Never stdout. The MCP spec (2026-07-28) requires that servers MUST NOT write anything to stdout that is not a valid MCP message — this was previously violated by 8 `console.log()` calls in engine captcha handlers, all replaced with `process.stderr.write()`.
**Tool manifest integrity (PR #27).** SHA-256 of serialized tool definitions (name + description) computed on startup. `createServer()` returns `{ server, manifest }`. If `EXPECTED_TOOL_HASH` env var is set, the server warns on mismatch and refuses to start if `STRICT_TOOL_HASH=1`. This catches tool definition drift between builds — a tool added, removed, or renamed is detected immediately.
**Engine router with fast mode (PR #28).** `fastMode: true` queries 5 high-value engines instead of all 13, reducing latency by ~60%. Engine tiers are derived from the frozen v1 benchmark:
Tier | Engines | Role
---|---|---
Core | Ahmia, Tor66 | High nDCG + high unique contribution
Reinforcement | Onionway, Excavator, NotEvil | Moderate nDCG, unique coverage
Low | 8 remaining | Minimal unique contribution
**Trust derivation from benchmark (PR #28).** Engine `trust` values are no longer hand-tuned. They are derived from frozen v1 per-engine nDCG and unique relevant contribution. The derivation is documented in
TRUST_DERIVATION_V1.md. The circularreasoningrisk(trustderived from the same SERPsused for replay)is mitigated by deriving oncefromv1, freezing, and not re-fitting.
**In-memory health history (PR #29).** Per-engine latency and status are retained in memory across searches, enabling trend analysis within a session. Not persisted to disk — consistent with the RAM-only design principle.
**SOCKS5 frame unit tests (PR #29).** Edge-case coverage for frame fragmentation across frame boundaries and unusual ATYP paths.
**Legal compliance framework (PR #30).** README now includes a legal compliance section. A private legal framework document (gitignored) covers jurisdictional analysis, acceptable use, and operational constraints. The project is licensed under PolyForm Noncommercial 1.0.0.
**Test framework migration (PR #33).** Migrated from `bun:test` to `vitest` for broader compatibility. 437 tests across 20 files, all passing.
**Brain module (PRs #34-36, migrated in #37).** An observation lifecycle module was built in three phases — observation/decay/session-state, consolidation/reflection/coherence, and conductivity/spreading-activation/hub-distillation. After the three phases, the brain module was extracted to a dedicated repository (`digital-brain-mcp`) to keep Ghost Search MCP focused on federated search. The migration was clean: the brain module is no longer in this repo, and the search ranking pipeline is unchanged.
### **Current state**
Metric | Value
---|---
Engines | 13 onion search engines, all with query contracts
MCP tools | 13
Tests | 437 (vitest, 20 files)
nDCG@10 (production replay) | 0.3944
nDCG@10 (benchmark evaluator) | 0.3900
Drift | +0.0044
License | PolyForm Noncommercial 1.0.0
LLM cascade | OpenRouter DeepSeek V4 Flash → NIM → Ollama → raw
Ranking profile | production-v1 (explicit, versioned)
### **What I did not do**
I did not re-open the benchmark. Frozen v1 stays frozen. The calibration cycle is closed. The replay adapter gives the signal to reopen if one ever appears — until then, the production decision (rrf-quality with RRF k=10, raw queries, trust-weighted) stands.
I did not add more judges, more rerankers, or another benchmark round. The system ordering (rrf-quality > rrf > bm25 > engine-count) is preserved in both the evaluator and production paths.
I did not pursue SOTA neural ranking methods (QUAM, REGENT, BlockRank). An external architecture review confirmed these are not applicable to federated snippet-based meta-search — they require full-text document access, which a meta-search engine does not have. The one viable SOTA path is LLM-based query expansion with RRF fusion (Exp4Fuse-style), which is what the query planner implements.
### **Remaining work**
* **Live Tor test with** `queryPlanner: true` — manual, operator-driven. The replay framework cannot measure this because sub-queries are generated dynamically. Only proceed with tuning if the live test shows a positive nDCG delta.
* **OpenRouter key rotation** — current API key has limited remaining credit.
* **Engine health trend persistence** — currently in-memory only. Cross-session persistence would require a storage layer, which conflicts with the RAM-only design principle. Under consideration.
### **Wiring (unchanged)**
benchmark v1 = frozen evidence + historical reference (eval-ranking.ts)
production-v1 = explicit executable ranking contract (profiles.ts)
replay = production SUT replayed against frozen v1 inputs (replay-production.ts)
intentional ranker change = run replay, compare delta, decide
engine failure != empty evidence (5-way classification)
retrieved content != trusted instructions
The production transition is complete. No benchmark phase 2 unless the replay gives a reason.