# STEP 1854 — Template-heavy domain family Phase 1a: characterization script + 8 synthetic samples + 5-axis measurement

**Timestamp**: 2026-09-07T00:46 (JST)
**Tab worktree**: main (rei-aios-97 [de6edb])
**Commit**: <commit 後に追記>
**Refs-STEP**: 1851

## 一行 summary

STEP 1851 design v0.1 の Phase 1a 実装。 characterization script (5 軸: bigram entropy / n-gram uniqueness decay / long-range MI / template ratio / compressor bpc gap) + 8 synthetic domain sample generator + 実測 CSV output。 **測定 pipeline は 動作 する** が Voynich signature v0.1 heuristic は 0/8 match = word-level を char-level に 意味論 誤転写 の honest finding、 v0.2 で 再校正 対象。 test 28/28 PASS。

## 主要 finding / evidence

### 実装
- `scripts/domain-family/characterize.ts` (5 axis 測定関数 + CLI runner + Voynich signature heuristic)
- `scripts/domain-family/generate-samples.ts` (deterministic PRNG seed 20260907、 8 synthetic domain)
- `data/domain-family/samples/` (8 files、 38 KB - 700 KB each)
- `data/domain-family/results/characterization-v0.1.csv` (実測 8 row × 12 column)
- `data/domain-family/README.md` (方法論 + 結果解釈 + honest scope + v0.2 送り)
- `test/step1854-domain-family-characterize-test.ts` (28 assertion、 各 axis の invariant + edge case)

### 実測結果 highlight
- **測定 pipeline は 動作**、 5 軸 全 domain で 数値出力、 直感 と 一致 の ranking:
  - bigram entropy: hash/uuid (4.05/4.04) > english (3.20) > regex (1.53)
  - uniqueness decay N: hash/uuid/numeric (6) < structured_log (12) < chess (15) < json/regex/english (21)
  - long-range MI: english (0.321) > json (0.174) > hash (0.001)
  - template ratio (n=3, top-10%): json (0.867) > log (0.796) > chess (0.581) > hash (0.129)
  - zstd bpc: english (0.20) < regex (0.33) < numeric (0.45) < json (0.69) < chess (1.55) < log (1.60) < hash (4.09) < uuid (4.15)
- **Voynich signature v0.1 heuristic (B≤5 AND |C|≤0.02 AND D>0.30): 0/8 match**

### honest finding — v0.1 heuristic の 意味論 誤り
藤本さん note の Voynich 主張 「4 語 で 決まり文句 消失」 は **word-level uniqueness decay**。 私 が Phase 1a script で 実装 した axis B は **character-level n-gram uniqueness decay** = 単位 が 違う。 word-level 転写 が axis B' として v0.2 で 必要。

## Honest scope

- 8 sample は **synthetic のみ** (real Voynich EVA / real English / real chess / real DNA 等 は Phase 1b、 外部 fetch 要)
- v0.1 heuristic threshold は **未 較正** (実 Voynich text 実測 後 に recalibrate)
- english_baseline zstd 0.20 bpc は **合成 text 過度圧縮** (5 paragraph × 20 repeat)、 実 English は Shannon bound 0.60-0.80 bpc (STEP 1847)
- axis E は **cmix 未含**、 zstd と gzip の gap のみ で proxy
- Voynich signature 定義 は **char-level vs word-level の 単位 差** で v0.1 は 意味論 誤、 v0.2 で 是正 必須

## Failure mode dataset entry

- **word-level → char-level の 意味論 誤転写**: 藤本さん note の 発見 単位 (word) を 私 が 実装 単位 (character) に 移し替えた が、 signature の 定量 主張 (「4 語 で 消える」) の 意味 は 単位 依存。 measurement pipeline は 動作 する が 主張 は 意味 不明 に なる。 v0.2 で word tokenizer + word n-gram を 追加 して 是正
- **synthetic sample の 過度圧縮 artifact**: english_baseline を 5 paragraph × 20 repeat で 合成 した ため zstd 0.20 bpc (実 English 0.60-0.80 bpc の 4x 過度圧縮)。 Phase 1b で 実 English (enwik8 fragment or arXiv abstract) に 差替 予定
- **top-p% ratio の 相対性**: templateRatio (top-10%) は unique n-gram の 数 に 依存 する 相対 量、 少ない unique 集合 では 「top-10%」 が floor で 1 個 に なり 意味不明。 test で 学び、 「相対的 な template 濃度」 として 解釈 する 注意 が 必要 (v0.2 で 絶対 threshold ベース に 追加軸 候補)

## Next STEP (藤本さん judgment)

1. **Phase 1a 承認** → Phase 1b (外部 fetch 6 domain: Voynich EVA / Esperanto Wikipedia / DNA / protein / real chess PGN / LaTeX) 着手 STEP 予約
2. **v0.2 axis 拡張** (word-level uniqueness decay 追加、 Voynich note 主張 直接測定) 別 STEP
3. **STEP 1852 Paper 180 concept と の 関係** (統合 vs 独立系統) の 再判断
4. **cmix 追加** (Windows binary / WSL 経由、 axis E の SOTA anchor)

**私 の 推奨**: Phase 1b (real Voynich EVA fetch + 追加 5 external domain) 先行。 word-level axis 追加 は Phase 1b と 同 STEP に 統合 (real Voynich text で immediate に validate)。

## 詳細参照

- 実装 files:
  - `scripts/domain-family/characterize.ts` (5 axis + CLI)
  - `scripts/domain-family/generate-samples.ts` (8 domain sample、 seed 20260907)
  - `test/step1854-domain-family-characterize-test.ts` (28 assertion)
- 出力 data:
  - `data/domain-family/samples/` (8 files)
  - `data/domain-family/results/characterization-v0.1.csv` (実測 CSV)
  - `data/domain-family/README.md` (詳細分析)
- 関連 STEP: STEP 1851 (design v0.1、 本 arc の 起点)、 STEP 1852 (対 arc Paper 180 concept)
- npm scripts: `test:step1854` / `domain-family:generate` / `domain-family:characterize`
