CodeEraser
Method

How it works

CodeEraser applies deterministic computation to the non-deterministic output of language models.

Why not a second model

A model writing into a long-lived repository stacks rather than edits: the same function twice, a fact restated in a third file, an update appended. Instead of auditing that drift with a second model, every verdict below is arithmetic over facts from the tree (token fingerprints, tree edit distances, graph in-degrees, git-window counts) against thresholds written down: no sampling temperature, no floating point in the judgment layer, no model in the loop. The same tree yields the same bytes on any machine at any hour, so a disagreement is settled by re-running the number and reading the file:line it points at.

Every heading links its full derivation in docs/reference/methodology.md, each constant traced to its implementing line.

How a verdict is made: Rust measures syntax units, token fingerprints, documentation shingles, git windows, the reference graph and a change's surfaces; Haskell judges structure and score, clones and same-role advice, documentation duplication, trajectory and audit, liveness and erase, and tombstone residue, one wire family per row; the gate and the per-family reports deliver the verdicts
One row per wire family: what Rust measures, the Haskell verdict it feeds, and what leaves (an exit code or report rows). Open the full-size SVG.

The nineteen judgment families

Three counts of “family”

Three counts ship across this site. nineteen booklets, one per judgment whose math earns a derivation: families 01–15 here, 16 on from the analysis track on a page of their own. sixteen judgment capabilities on the wire (the stack diagram's count, beside the handshake hello, the definition package tables/1 and the report documents document/1; the join and the split advisory ride inside verdict and structure requests, and the FPR tier ladder is a discipline, not a request). twenty-one read-only MCP tools follow the report surfaces, not the wire: churn and graph --sites, erase (its dry-run plan only), doctor (machine state), erase_log (the applied-erase trail), update_check (this build against the latest release), similar_units (ce similar's advisory document over similar/1), and query and rules over query/1 (its own page).

01T1/T2 clone detection — winnowing fingerprint index

Finds exact and parameterized duplicate token runs: renamed variables and changed constants are clones, changed syntax is not.

t = window + kgram - 1 = 26 + 25 - 1 = 50 tokens

Any common run of at least t normalized tokens holds t - k + 1 = w consecutive k-grams, one complete window, whose minimum depends only on its contents, so both copies select it: at least one shared fingerprint, always, the Schleimer et al. SIGMOD'03 no-miss bound held as a correctness contract. The rolling hash is Rabin-Karp, h = (h - t[i-k]*top)*BASE + t[i], over FNV-1a leaf hashes.

Derivation and constants
  • kgram 25
  • window 26
  • guarantee t 50
  • BASE 1_000_003
  • HOT_CAP 64
  • DEFAULT_MIN_DISTINCT 7
  • near-miss band 25 <= len < 50
  • TOKENIZER_REV 4
  • schema ce.dedup-report/0.5.0

02T3 near-miss clones — tree edit distance (TSED)

Judges units whose ASTs are almost the same but whose token streams are not: reshaped and rewritten copies.

TSED(a, b) = (max(n1, n2) - ted(a, b)) / max(n1, n2)
clone      ⇔ TSED >= 0.85
cloneDecidesWith (num, den) t n1 n2 = (mx - t) * den >= num * mx  where mx = max n1 n2

ted is Zhang-Shasha with unit costs (delete = insert = 1, relabel = 0 on matching kind codes, else 1), always integral, so the comparison is an exact cross-multiplication decidable both ways: at max = 100, ted 15 is a clone and ted 16 is not. Two O(1) prefilters, q · tsedDen < tsedNum · max for q ∈ {min(n1,n2), I} with I = Σ_label min(c1, c2), cut pairs before any TED runs; they are admissible because they use the same 85/100.

Derivation and constants

Since 1.8.0 verified pairs ride ce check's clone axis and ce join's similarity leg as sim kind 1, each verdict cached under the two trees' content keys until the proto or a knob changes.

  • tsedNum 85
  • tsedDen 100
  • T3_MIN_NODES 24
  • unitNodeCap 256
  • pairCap 4096
  • LSH_SHAPE (128, 32, 4)
  • HOT_GROUP_CAP 64
  • schema ce.clone-report/0.3.0

03Documentation duplication — shingling + MinHash/LSH

Decides whether two documentation blocks (markdown paragraphs, HTML text, plain-text paragraphs, comment blocks, docstrings) are near-duplicates.

dupDecidesWith num den inter union  =  inter * den >= num * union

dupVerdictWith (num, den, vfloor) inter union run
  =  dupDecidesWith num den inter union  ||  run >= vfloor

Bound to (80, 100, 50): Jaccard ≥ 0.80 decided in integers, or a verbatim run of at least 50 words. The core computes inter and union itself from the ascending deduped shingle sets, so raw counts cross the wire, never a ratio. MinHash/LSH is an RNG-free coarse filter (the permutation index is the salt); a run of R shingles spans R + k - 1 words.

Derivation and constants
  • jaccardNum 80
  • jaccardDen 100
  • shingleK 5
  • minDocTokens 50
  • verbatimFloor 50
  • docLineCap 200
  • licHeadLines 5
  • DOC_SET_CAP 8192
  • DOC_PAIR_CAP 4096
  • DOCDUP_REV 8

04Structure judgment — tree-scale entropy, eight axes

Judges the tree rather than the file: directory geometry, naming distributions, reference locality, documentation coverage, doc staleness, redundancy, modularity.

tsallis2 cs     = 1 - Σ (c/N)^2            -- and = 0 when N == 0
tsallis2Norm cs = tsallis2 cs / (1 - 1/n)  -- n = nonzero bins, n > 1
chi2 pairs      = Σ_{r > 0} (p - q)^2 / q  -- p = o/Σo, q = r/Σr
perMille r      = floor (r * 1000)
charge_i = floor(scale * v_i / (v_i + N))   -- v = flagged dirs, N = dir total
raw      = Σ_axes (charge_i * violCost)
score    = max 0 (scale - raw `div` (structViolCostNeutral * judgedAxisCount))

Shannon entropy and KL divergence need logarithms, not exactly decidable, so the family uses their rational-closed relatives, Tsallis-2 diversity and the χ² f-divergence, over Data.Ratio. Observed mass on a zero-reference bin makes χ² Nothing, a refusal that names those directories. An axis's mass v is its flagged-directory count, charged into ‰ of scale and folded through structViolCostNeutral = 10; judgedAxisCount is 5 to 8, by which optional fact tables rode the wire.

Derivation and constants
  • violCost 10
  • scale 1000
  • depthCeil 8
  • fanoutCeil 30
  • namingMin 5
  • namingCeil 600‰
  • mixRefFloor 5
  • misplaceMin 3
  • bigDirFloor 8
  • staleMin 1
  • dupMin 1
  • deadMin 1
  • structNodeCap 524288
  • modFloor 1
  • modMassFloor 4

05Scoring and the ADR-006 ratchet

Folds seven axes into one 0–1000 score and gates it against a banked baseline that may only tighten.

raw     = sum_i (w_i * p_i * violCost)
wTotal  = sum_i w_i                            -- derived, never a literal
score   = max 0 (scoreScale - raw `div` (violCostNeutral * wTotal))
p(x) = 0                              if x <= S
     = pMax                           if H <= S          -- degenerate fallback
     = pMax * ((x - S) / (H - S))^2   if S < x <= H
     = pMax * (1 + 2*(x - H)/(H - S)) if x > H           -- C¹ linear arm

m = median(x)
r = median( max(x/m, m/x) )                    -- >= 1 by construction
S = clamp(floor(m * r^k), [softMin, softMax])
tolerated(c) = max (c * tolNum `div` tolDen) (c + tolAbs)

added   = current \ baseline        -- non-empty => fail
removed = baseline \ current        -- informational; drives the shrink

Axis 0, the only one that is not a count, is a convex penalty on file size in exact Rational, monotone past the hard line; pMax is the value at H, not a ceiling. The soft line S is the repository's own frozen LOC statistic, S = clamp(median + k·MAD, …) rewritten multiplicatively so no logarithm is taken, derived at establish and frozen into the baseline. The fail bit is a disjunction of six named conditions, ratchet_over, discrete_added, floor, dedup_budget, knobs_digest, rows_dropped, printed after FAIL in that order.

Derivation and constants

Since proto 3.1.0 a continuous row may carry its path class, the 1-based index of the first matching [[rules.class]] (0 for none), and is charged against that class's soft line, hard line and CoC ceiling, the global line as fallback. Beside the ceilings' codes 0 / 1 / 2, classKnobs adds 3 (proto 5.1.0), the class's ratchet allowance in lines, and 4 (proto 6.4.0), its cognitive-complexity allowance; declared, either replaces both global legs, so 0 means any growth is over. The baseline stays three columns; names and globs never cross the wire.

  • scoreScale 1000
  • violCost 10
  • defaultWeight 1
  • sizeHard H 750
  • sizePMax 10
  • softLineK 2
  • [softMin, softMax] [200, 500]
  • sizeCeil fallback 300
  • cocCeil 15
  • deadIndegCeil 0
  • rewriteNum/Den 50/100
  • sccFloor 2
  • tolNum/tolDen 102/100
  • tolAbs 10
  • legs cross at 500
  • verdictNodeCap 131072
  • verdictRowCap 524288

06Graph liveness and dead-code verdicts

Answers "which files does nothing live reach?" over a reference graph whose node identity is the row index, with no text on the wire.

arcs  = { (s,d) | [s,d,kind,rung] ∈ edges,  rung <= minRung,  kind ∉ inert }
reach = ⋃ { reachable(G, s) | s ∈ entries(entryMask, flags) }
public     = testBit flags 0
referenced = indeg >= 1 over kept arcs
judged     = i ∉ reach
code       = 1 + public + 2*referenced    -- the lookup table is the authority

Four codes keep an exported-but-unreferenced API from collapsing into plain dead: 1 unref_private, 2 unref_public, 3 unreach_private, 4 unreach_public. A site walks its language's rungs in order and the first rung yielding exactly one in-scope candidate wins; more is Unresolved, and External is a correct answer, not a miss. Cycles are reported, never judged; a cyclic island with no entry seed is dead by reachability alone.

Derivation and constants

inert = {3 asset, 5 unused ref-def}, excluded by the core: an image link renders bytes and an unused reference definition renders nothing. Since proto 2.32.0 each dead row carries a confidence judged from the per-language site ledger (0 unvouched, 1 vacuous, 2 vouched), the erase family's trust boundary. Import-edge precision: 38/40 = 0.95 over five pinned corpora against a ≥ 0.90 gate, on a sample frozen before any resolver existed. The six languages plan v2.30 added each passed a frozen, blind-audited exam: C 73/73, C++ 15/15, Java 36/36, Lua 86/86, R 18/18, HTML 75/75 (recall 73/74, 15/17, 36/36, 86/86, 18/18, 75/76), and a per-language false-positive ledger (six corpora, 400 commits each) shows 0 false intercepts; see EVAL-SET-LANGS and FPR-REPLAY.

  • minRung 5
  • entryMask 126 (bits 1–6)
  • sccFloor 2
  • nodeCap 131072
  • edgeCap 524288
  • GRAPH_REV 23
  • tsconfig extends unbounded, cycle-checked
  • precision gate ≥ 0.90

07The three-signal join

Combines similarity × graph position × churn into one of four candidate codes, then deliberately declines to act on it.

(1, sev 2, [1,2,3,4], [])    -- merge_candidate:  sim + graph + both referenced + distinct SCCs
(2, sev 3, [1,2,5],   [6])   -- delete_candidate: sim + graph + dead flank, RG10 guard clear
(3, sev 1, [1,2,7,8], [])    -- churn_hotspot:    sim + graph + cochange + rewrite
rewriteHot  = total > 0 && (rewrote_a + rewrote_b) * rewriteDen >= total * rewriteNum
cochangeHot = cochange >= cochangeFloor
legsMask    = legSim .|. (if graphBoth then legGraph else 0) .|. legChurn

Priority is data, not guard order: the first row whose required bits hold and forbidden bits stay clear wins, else code 0, so the property battery can falsify the order by rotating the table. Every gating row requires the graph bit: mask 5 (graph leg absent) carries only code 0 rather than pretending indegree 0.

Derivation and constants

Since proto 2.33.0 each candidate row carries a leg-agreement confidence, the severity face ships once as joinSeverity, and ce join judges over the same verdict/1 road as the check gate. No join code enters the fail bit: ce join exits 0 unless the run fails, and a missing or refusing core is a run failure (exit 2).

  • rewriteNum/Den 50/100
  • cochangeFloor 2
  • entryMask 126
  • legSim 1
  • legGraph 2
  • legChurn 4
  • cochangeFileCap 20
  • schema ce.join-report/0.4.0

08Split-ROI seam pricing (four legs)

Prices "is this long file worth splitting, and where?" as a number, an advisory outside the structure score and the fail bit.

benefitMilli(u) = max 0 (floor (1000 * (p(total) - p(end_u) - p(total - end_u))))

costMilli(u)    = crossRefs(u)      * roiRefMilli
                + cutClones(end_u)  * roiCloneMilli
                + crossChurn(u)     * roiChurnMilli
                + roiPhiMilli

viable          ⇔ b >= c            -- ROI >= 1, evaluated without division

Benefit is the graded-zone penalty a split returns, on the same convex curve the verdict family uses (imported, not re-derived); p is convex with p(0) = 0, hence superadditive, so the bracket is non-negative. The best seam is the exact rational argmax over ROI, compared by cross-multiplied b % c; n top-level units give n - 1 seams. Long-and-cohesive gets an exemption with numbers attached, long-and-splittable a cut line.

Derivation and constants
  • seamSoft S 300
  • seamHard H 750
  • seamPMax 10
  • roiRefMilli 250
  • roiCloneMilli 500
  • roiChurnMilli 150
  • roiPhiMilli 500
  • NAME_FLOOR 3
  • CHURN_WINDOW_DAYS 14
  • knob codes 12–18

09Edit four-classification (update supervision)

Reduces every supervised edit to four integer counts per file pair (matched, novel, moved, deleted), so "stacked, not applied" becomes a measurement.

siteOpens s n  =  n * movedCost + s  <  n * plainCost

destFloor      =  least n with siteOpens siteCostCross n   =  2
accepted   = isStart && distinctEvidence >= destFloor && anchored
anchored   = any (\(_,_,w) -> w >= anchorFloor) evidence

The cross-file evidence floor is derived: siteCostCross = 2 makes a single cross line a tie (1*1 + 2 = 3 = 1*3), ties do not open, so destFloor is 2, and that tie is the coincidence rejection. Only anchorFloor = 19 is decided: in the dual-corpus shadow ablation the invented station's widest anchor measured 16 alnum characters and the thinnest real anchor 19, so 19 tops the window that kills every measured coincidence and keeps every real site. The L2 delta moves one way: plain lines may become moved, never the reverse.

Derivation and constants

Since proto 7.1.0 paired declRem / declAdd tables of declaration keys that vanish on one side and appear on exactly one other open an edge when the spans share declFloor = 1 leftover hash (the key pays declCredit = siteCostCross). They append relocations with lines = 0; no line class, block, suspicion or score moves.

  • movedCost 1
  • plainCost 3
  • siteCostWithin 0
  • siteCostCross 2
  • destFloor 2
  • anchorFloor 19
  • MAX_D 3000
  • MAX_BRIDGE 7
  • bucketCap 64
  • stackingNovelFloor 20
  • stackingRatio 10

10Score trajectory — the trend slope verdict

Is the check score going up, flat or down across history, and by how much per day?

x_i   = ts_i % 86400                 -- seconds to days, exact ratio
y_i   = (score_i * 1000000) % scale_i

slope = median{ (y_j - y_i) / (x_j - x_i) : x_i ≠ x_j }   -- Theil-Sen
slope < -band  → 2  (degrading)
slope >  band  → 0  (improving)
otherwise      → 1  (flat)        where band = floorMicro

% is Data.Ratio's exact-ratio constructor, not modulo; y puts every commit on a fixed 10⁶ full-scale grid, so rows under different scoreScale values compare. The slope is the median of pairwise slopes (trend/2): one wild point drags a least-squares mean anywhere and cannot move the median past its neighbors.

Derivation and constants

The judged view sorts by timestamp and keeps the tsWindow most recent points, so rebased commits are legal. The steepest single-step fall ships as cliff, the longest strictly-falling run as declineRun, each by request index; hashes never cross. Below minPoints, or with no timestamp-distinct pair, the slope is Nothing, never a fabricated flat; at the default floor 0, degrading is reported and cannot fail.

  • minPoints 3 (knob 0)
  • declineFloorMicro 0 (knob 1)
  • row + knob cap 4096
  • tsWindow 512
  • full-scale grid 10⁶
  • schema ce.trend-report/0.3.0

The FPR discipline: what earns a rule the right to deny

11FPR discipline and the guard tier ladder

A deterministic gate over a non-deterministic writer: on PreToolUse for Write|Edit the guard answers each pending write with exact arithmetic, and a rule class enforces only once it has paid for the right.

TIERS = ["observe", "warn", "ask", "deny"]
PROMOTED_DEFAULT = "deny"
cap      = lines_for(file).file_lines_fail     // the file's class table; 750 for class 0
breach   ⇔ cap != 0 && lines > cap

permille = (lines - S) * 1000 / (H - S)
0   ..= 249  →  observe
250 ..= 750  →  warn
751 ..       →  ask
Derivation and constants
tierpermissionDecisioneffect
observenone, returns before printingfeed line only, no injected text
warnallowedit proceeds, reason surfaces as a visible warning
askaskthe user is prompted
denydenywrite is refused, reason points at the existing file:line

Determinism comes from replaying the write: resulting_lines computes the exact post-write line count, and a call that would fail on its own (missing file, ambiguous non-replace_all match) returns None, so the rule stays silent. An unrecognized [guard] mode resolves to observe (ce.toml ERROR: …), never passed through.

A class earns deny by its record. M3: ≤ 1 mis-block in 500 real normal edits, N=1 demonstrations disallowed. M4: FPR ≤ 1% over 500 real normal edits on a set pre-registered before implementation, ≥ 200 edit samples, ≥ 50% from real agent transcripts, sampled only from observe-mode and pre-guard sessions so the guard cannot certify itself. Two classes have paid: T1/T2 exact duplicate write, and hard-budget breach at file > its hard line (750 by default, or the file's [[rules.class]] line), so the hook denies exactly where the CI wall fails. Every other rule stays at observe for want of its own record.

Since CodeEraser 1.2.0 the duplicate-write rule charges novel duplication only: the K-round replay of 2,761 real edit events (this repository's history plus the pinned requests tail) separated genuine landings from re-fires on already-budgeted files and split/fold mid-states, so matches the replaced content already carried (Write: the on-disk file; Edit: old_string) are subtracted. A rewrite carrying its debt is silent; an introduction, or a split to a new file, denies and teaches the ordering that passes (trim the source first). Both framings: FPR-REPLAY.md.

Replaying linear git history as an edit stream at the shipped knobs t = 50 and min_distinct = 7: 630 events, 35 blocks, 0 false after arbitration → 0.00 per 500 (M3), the 35 blocks remediated as the clone-block budget stepped 251 → 211 → 209 → 205 → 202. The standing instrument (cli/tests/it/fpr_replay.rs, 2026-09-05): requests 365 events / 0 intercepts; 3164 self-repo events with 178 landed introductions and 79 write-first mid-states.

The graded zone reads S and H off the hard budget's per-file table. [guard] zone_tiers is off by default: its pre-registered rule armed only if every corpus read ≤ 1 %, and the first run read self 0.54 % but requests 2.46 % (variant B), held by cli/tests/it/fpr_zone_gate.rs. Warns are rate-limited to once per (rule, file, session) and clipped at a token budget; a deny is not.

  • file_lines_warn S 300
  • file_lines_fail H 750
  • zone_tiers false
  • FPR gate ≤ 1% / 500 edits
  • WARN_BUDGET_TOKENS 200
  • STOP_BUDGET_TOKENS 400
  • CHARS_PER_TOKEN 4
  • feed schema ce.observe/0.13.0

Acting on the verdicts: erasure and advice

12Deterministic erase — the safety predicate

Judges erase-plan rows for three provable classes (dead_file, verbatim_doc, t1_twin) and refuses every unsafe row with a named reason code.

class  = 0 (retired 4.0.0, refused by name) | 1 verbatim_doc | 2 t1_twin
         3 dead_file (confidence road, 2.32.0)
reason = 0 eraseable | 1 language_unresolved | 2 not_full_segment
         3 bytes_differ | 4 copy_not_dead | 5 unit_not_covered
         6 public_surface (6.1.0)

Rust assembles integer facts from the three source families; Haskell applies a fixed first-failure predicate. Since proto 2.32.0 class 3 trusts the graph family's per-row confidence; since proto 6.1.0 a public dead verdict (2 unref_public, 4 unreach_public) is refused as public_surface before confidence is weighed, so an exported API is never eraseable; only then does an unvouched confidence refuse the row.

Derivation and constants

Class 0, the local-count road, retired at proto 4.0.0; the core refuses it by name. A degraded over-cap reply authorizes nothing, and the safety surface has no knobs.

  • classes 4
  • reason codes 7
  • eraseRowCap 4096
  • since proto 2.16.0
  • wire erase/1

13Unmentioned-declaration advisory — the mention veto

Lists every judged declaration whose name no other file spells, with a visibility code: a negative instrument that never judges.

veto  = another file spells it | fold (Rust, ≥2 segments ∧ ≥7 chars)
        | the file's own exception regions spell it
row   = [node, vis, conv]      (integers only — a name never rides the wire)
emit  ⇔ vis ⊇ {exported, scope-exported} ∧ conv has no exempt bit (0..10)
code  = 1 private > 2 restricted > 3 reexported > 0 public
Derivation and constants

U is its own walk: hidden files enter, .git/.ce are cut by name, an undeclared nested repository is cut whole, and a path the root's .gitmodules declares stays as foreign (it can spell a name, never measured). Only .gitignore and .ceignore are honoured, so one commit yields one U, and every walk parameter is a MENTION_REV input; the store keeps 64-bit hashes, never text. Haskell reads only the visibility word and the mounts row (private mod mounts, façade re-export, package-private). Both tables ride graph.request together or not at all; without them the reply is the ten-key one byte for byte and the dead set is unchanged. Past unmentionedCap the core drops the table and says so. Never a gate, never in ce erase.

  • MENTION_REV 4
  • FOLD_MIN_CHARS 7
  • unmentionedVisMask 3
  • unmentionedCap 131072
  • unmentionedHardCap 524288
  • since proto 6.2.0
  • wire graph/1
Three lanes (measured facts in Rust, judgment computations in Haskell with their formulas, and the gates that decide the exit code) with the constants the sections above cite written on the diagram
The constant sheet: Rust extracts facts, Haskell turns them into verdicts with the formulas shown, the gates decide the exit code; every number on it is cited above. Open the full-size SVG.

The change-time family: what an erasure leaves behind

The families above read a tree; the fourteenth reads a change (one edit, session or commit) at the PreToolUse hook, the Stop audit and ce precommit / ce commitmsg.

14Tombstone residue — the erased-name conjunction

Catches the change that erases X from the code and writes it back as an absence: a heading (no X), an identifier without_x, a sentence X is no longer needed.

R     = ⋃ names(before) \ ⋃ alive(after)      (structural positions only; keys, never text)
row   = [kind, marks, names]                  kind ∈ {0 bracketed, 1 bare, 2 prose}
site  ⇔ names ≥ 1 ∧ (kind ≠ 2 ∨ marks ≥ 1)    (the conjunction, read per sentence)
over  ⇔ budget declared ∧ |sites| > budget     (spoken at [tombstone] tier; ships observe)
Derivation and constants

A name is what a text declares: an identifier outside comments and literals, a unit name, a heading, a list lead; an inline code span only mentions, keeping a name alive. Rust sends three integers per added surface; Haskell seats the sites and judges the budget. A changelog-role document is exempt whole, by path convention, ledger shape or [tombstone] ledger; a quote run or section that is itself a ledger exempts only itself. The class ships at observe; promotion would argue from the nine-round FPR replay over two histories in docs/FPR-TOMBSTONE.md.

  • TOMBSTONE_REV 1
  • min_ascii_name 3
  • min_wide_name 2
  • join_max 3
  • SEGMENT_TOKENS 3
  • minName 1
  • minMarks 1
  • tombstoneRowCap 65536
  • HASH_CAP 256
  • SITE_CAP 10
  • since proto 6.6.0
  • wire tombstone/1

The advisor family: the units that play the same part

Clone families need two units to look alike; the fifteenth finds units that play the same part with no line in common, as advice at ce similar, the MCP tool similar_units, the GUI's similar screen and one line of the Stop audit.

15Same-role advisor — sparse retrieval and in-repo association

Ranks the units most like one unit (or a text) by integer BM25 over six-channel term bags from the index's facts, and asks the core which play the query's role: an offline "code RAG", no model, no float, advisory only.

bag    = six channels N P C D S L        (facts the parse already carries; hashes, never words)
score  = Σ w · idf · 22·tf·avg / (10·tf·avg + 3·avg + 9·len)   (k1 = 6/5, b = 3/4; integer fixed point)
widen  = top-m PPMI neighbours at ≤ ½ weight    (opt-in view; never evidence)
role   ⇔ (N ≥ 1 ∧ C ≥ 1) ∨ (N ≥ 2 ∧ shapeEqual)  (judged in Haskell over similar/1; advisory only)
Derivation and constants

Each T3-universe unit gets one bag (name pieces, shape, callee spellings, its doc segment, node-kind histogram, literal kinds), every word stemmed and hashed with its channel letter. Bags persist in the content-hash-gated refresh (bag is its own posting list; a pair table measured 7–10× the cold index and was not built, so co-occurrence is derived at query time). The in-memory instruments and the SQL reader share one trait-written ranking and agree on every unit of five corpora. Rust sends the query bag and a nine-integer row per candidate; Haskell orders exact rationals and applies the role conjunction. Floors from two arbitrated oracle generations: 60 % on the first sample, 40 % on the holdout, which retired every tuning candidate the first sample favoured. Never a gate, never in ce erase.

  • SIMILAR_REV 1
  • K 5
  • W_UNIT 256
  • TERM_CAP 96
  • TOP_M 3
  • MIN_COOC 2
  • roleMinName 1
  • roleMinCallee 1
  • roleMinNameShape 2
  • similarCap 65536
  • since proto 6.7.0
  • wire similar/1

Two honesty boundaries

PreToolUse shapes behavior; it is not a security boundary: Bash: echo >> or sed -i bypasses it, and the backstops are the Stop audit over git diff (write-tool agnostic) and the CI gate. The hook is deliberately fail-open: an internal failure allows the edit and lands in the observe feed instead of reading as "no duplicates". Both are written into the plan: a gate that lies about its reach is worse than one that states it.