Files
Arcrun/kbdb/wrangler.toml
T
uncle6me-web b9b5d98d75 併 bge-m3 換代(#59)+補上分支漏做的另一半:Vectorize index 也換
leo 2026-08-05:「換 embed model 當然要合併,當然要換 vectorize,**原本的根本不能用**」
——我把已拍板的事當成待裁決,錯了。

## 為什麼非換不可(08-03 leo 拍板,5 組中文測資實證)
  @cf/baai/bge-base-en-v1.5   768   排序正確 2/5   margin -0.0413   1660ms  ← 舊
  @cf/baai/bge-m3            1024        5/5        +0.1410    959ms  ← 新
舊英文模型嵌中文:分數全擠在 0.65-0.81,區辨力接近沒有=**根本不能用**。

## 🔴 分支只做了一半,另一半我補上
feat/embed-model-m3-t59 只改 kbdb(embed.ts / types.ts / 測試),**完全沒動 deploy.ts**
⇒ 就算合併,安裝器仍會建 **768 維的舊 index**,新向量根本收不進去。
(git show f3af5d3 --stat 實查:只有 3 個檔,全在 kbdb/)

本次補上安裝器那半:
· KBDB_VECTORIZE_INDEX: arcrun-kbdb-embed → **arcrun-kbdb-embed-m3**(換名字,非只換維度)
· ensureVectorizeIndex: dimensions 768 → **1024**
· kbdb/wrangler.toml 說明與 metadata-index 指令同步改新名(照抄不會建錯 index)
· embed.ts docstring「768 維向量」→ 1024(分支漏改,會誤導)

## 為什麼要換「名字」不是只改維度
① 維度 768→1024,舊 index 收不進新向量
② 就算維度相同也不能沿用——不同模型的向量混在同一 index 比對出來是垃圾,
   而 #58(Vectorize vector delete 未接)代表舊向量刪不掉
   ⇒ **開新名字反而順手繞開 #58**,且新舊並存可回滾。

## 查證(歷史警察 Q0-Q3,非假設)
· git log --all -S"arcrun-kbdb-embed-m3" / -S"dimensions: 1024" 皆空 ⇒ 無分支改過,非重造輪子
· **metadata index 用同一個常數**(deploy.ts:426 ensureVectorizeMetadataIndexes)
  ⇒ t36 的四個 metadata index(owner_id/entry_type/source/library,Arcrun#11 根因修復)
  會自動建在新 index 上,**不會因改名而遺失**——這點特地查過,不是假設。

## 驗
· kbdb 測試 87/87 綠(含新增 4 項 embed-model-config)
· cli tsc --noEmit **零錯誤**
· 三處一致性機械確認:embed.ts=bge-m3/deploy.ts=dimensions 1024/index=arcrun-kbdb-embed-m3
· 踩到並記錄:註解寫 `**dimensions=1024**​/metric` 會因 `*/` 提早關掉 block comment(TS1127)

## 既有實例遷移(尚未執行,需對實例操作)
建新 index → 重部署 kbdb(binding 指新 index)
→ POST /embed/backfill {"reindex":true} 到 remaining=0 → 舊 index 可刪。

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-05 13:46:33 +08:00

62 lines
4.3 KiB
TOML
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
name = "arcrun-kbdb"
main = "src/index.ts"
compatibility_date = "2025-02-19"
workers_dev = true
compatibility_flags = ["nodejs_compat"]
# KBDB Base — atomic universal table (SDD .agents/specs/arcrun/kbdb-base).
# Base needs D1 ONLY (free, no credit card). embed module adds Vectorize+AI bindings
# (optional, self-host opens it themselves). triplet is a separate repo.
[[d1_databases]]
binding = "DB"
database_name = "arcrun-kbdb"
database_id = "0c580910-e00b-4f8e-9c57-ac54ea52242f" # 官方 prod D1arcrun-kbdb);self-hosted deploy.ts 會注入用戶自己的 id 覆蓋
[vars]
ENVIRONMENT = "production"
# ── Auth guard (t115 二修, fail-closed) ────────────────────────────────────────
# The installer generates a random token at deploy time and secrets it into BOTH workers:
# wrangler secret put KBDB_INTERNAL_TOKEN (arcrun-kbdb)
# wrangler secret put KBDB_INTERNAL_TOKEN (arcrun-cypher-executor)
# cypher sends the token as `Authorization: Bearer <token>` via kbdbBase().
# Workflow http_request nodes that hit KBDB directly must include
# `Authorization: Bearer __KBDB_TOKEN__` (installer substitutes the value).
#
# Secret NOT set → writes (POST/PATCH/DELETE) are rejected 401 immediately (fail-closed).
# Reads (GET) pass with a server-side warning — old instances survive the upgrade
# window until both workers receive the secret at the same time.
# Secret SET → all non-health routes require correct Bearer; / and /health exempt.
# ──────────────────────────────────────────────────────────────────────────────
# ── Optional embed module (issue #7 / SDD T2.4) ────────────────────────────────
# Base 預設不開(free-tier 友善)。self-host 開語義查詢時,deploy.ts 偵測 config kbdb_embed:true
# → 取消下面兩段註解(注入 active binding)並 `wrangler vectorize create arcrun-kbdb-embed-m3
# --dimensions=1024 --metric=cosine`**bge-m3 = 1024 維**)。官方帳號同理由 deploy 注入。
# 🔴 2026-08-03 換代(leo 拍板,5 組中文測資實證:舊英文模型排序 2/5、margin −0.0413=根本不能用;
# bge-m3 5/5、+0.1410、959ms):`bge-base-en-v1.5`(768) → `bge-m3`(1024)。
# **換模型必須換 index**:① 維度 768→1024,舊 index 收不進新向量
# ② 就算維度相同也不能沿用——不同模型的向量混在同一 index 比對出來是垃圾,
# 而 #58Vectorize vector delete 未接)代表舊向量刪不掉 ⇒ 開新 index 反而順手繞開 #58。
# 既有實例遷移:建新 index → 重部署 kbdbbinding 指新 index
# → `POST /embed/backfill {"reindex":true}` 重嵌到 remaining=0 → 舊 index 可刪。
# ⚠️ Arcrun#11:光建 index 不夠。要對 owner_id/entry_type/source 下 filterowner-scoped/類型-scoped 語意查詢),
# 必須另建 metadata index,否則帶過濾一律回 0 命中:
# wrangler vectorize create-metadata-index arcrun-kbdb-embed-m3 --property-name owner_id --type string
# wrangler vectorize create-metadata-index arcrun-kbdb-embed-m3 --property-name entry_type --type string
# wrangler vectorize create-metadata-index arcrun-kbdb-embed-m3 --property-name source --type string
# wrangler vectorize create-metadata-index arcrun-kbdb-embed-m3 --property-name library --type string
# libraryportal-auth P1「庫」filterupsert 端把未標記正規化成 'general',查詢走 $in
# metadata index 只收「建立後 upsert」的向量 → 既有向量須 `POST /embed/backfill {"reindex":true}` 重推
# (建 library index 後同樣要 reindex,否則舊向量帶 library filter 一律 0 命中)。
# deploy.ts 的 ensureVectorizeMetadataIndexes() 已把前三個 index 隨部署冪等建好;library 待補進該清單
# cli/ 屬 portal-auth P1 範圍外,見 portal-auth tasks.md 部署清單附註)。
# 沒有這兩個 binding 時,kbdb/src/embed.ts 的 embedEnabled() 回 false → 維持 LIKE keyword、API 不變。
#
# [[vectorize]]
# binding = "VECTORIZE"
# index_name = "arcrun-kbdb-embed-m3" # bge-m3 1024d2026-08-03 換代)
#
# [ai]
# binding = "AI"