Kush V9 Is Held From Production
Kush V9 loaded successfully on the self-hosted Mac Mini runtime, but load success is not the same as release readiness. After Prime Agent E2E comparison and direct live prompts, we rolled production back to hotep-llm-kush-v82.
The candidate did fix one important serving problem: it ran with the native Llama 3.1 chat template instead of a forced ChatML template. That made the adapter behavior more visible and gave us a clean test boundary.
Prime Agent E2E Comparison
The direct comparison used seven representative prompts against the validated Kush V82 baseline and the Kush V9 candidate.
| Model | Passes | Entity passes | Avg persona | Avg latency | Avg tok/s | CJK | Rubric | Repetition |
|---|---|---|---|---|---|---|---|---|
| hotep-llm-kush-v82 | 4/7 | 4/7 | 21.96 | 8.610s | 22.88 | 0 | 0 | 0 |
| hotep-llm-kush-v9 | 4/7 | 4/7 | 19.08 | 14.343s | 6.14 | 0 | 0 | 0 |
V9 did not beat V82. It tied the small pass count, scored lower on persona, and ran materially slower.
Why It Was Pulled
The live candidate still missed hard entity recall:
- Strong Dad must resolve to Trey Quinn.
- Uncle Hotep must resolve to Demond Handy.
- Grifties ticket prompts must point to grifties.com instead of inventing free or informal access.
- Shakka Ahmose needs grounded classical African history coverage, not generic public-figure hallucination.
Those misses are concrete enough to fix. They also prove the promotion gate was too weak: a model that merely loads and produces clean English can still be wrong on community-critical facts.
What Changes Next
V9.1 will be trained and promoted only after it clears a hard entity gate with required and forbidden terms for the identities that failed here. The release path now treats Prime Agent E2E, direct live prompts, entity recall, speed, and runtime health as separate gates.
Hotep Intelligence remains sovereign infrastructure: local model, local serving, local evidence, and public correction when the evidence says a candidate is not ready.