The Gauge Is the Obstacle
Task Identity, Cross-Family Transfer, and Merge-Conflict Prediction from Adapter Weights Alone
Abstract
We train a pre-registered bank of 480 LoRA adapters—2 model families × 6 tasks × 40 seeds (Qwen2.5-1.5B-Instruct and Llama-3.2-1B-Instruct)—and run five locked weight-only analyses against it exactly once, under a completeness interlock, with every open analytical choice pinned in a dated ruling before the analysis that consumed it ran—all but one by an independent Director before the bank completed; the exception, the D3 pair design, was declared minutes after completion, before any pairs, merges, or labels existed, and ratified by the Director on review. (1) Raw adapter weights do not reveal their training task (leave-one-out accuracy 0.0792 qwen / 0.1375 llama, both failing the 1.5×-chance lock at chance 0.1667), while the GL(r)-gauge-canonical and vocab-signature representations identify all six tasks at 1.0000, permutation p = 0.000999. (2) Our pre-registered prediction that task structure would not transfer across families is refuted—and refuted specifically by the triviality control added in review: raw transfer-at-chance was a family-scale artifact (family-identity probe 1.0000 on raw features), and shift-controlled transfer runs 0.7375–0.7833 (binomial p ≤ 1.20×10−84). (3) The identity backbone is the load-bearing structure: swapping trained bridges across tasks costs less than +0.001 nats of validation loss; only full-entry permutation that destroys the backbone costs +2.8086 (qwen) / +3.8365 (llama). (4) Weight-only features predict midpoint-merge conflict at group-aware AUC 0.995 [0.983, 1.000] (qwen) / 0.962 [0.898, 0.999] (llama), a +0.320 / +0.249 margin over a distance-only baseline with both margin CIs excluding 0—in the weight-only, post-hoc regime, distinct from the training-time form of merge-conflict prediction (Tang et al. 2026). (5) A pilot bridge-deviation ↔ generalization-gap correlation of r = 0.888 shrinks to r = 0.300 [0.175, 0.415] pooled at bank scale—real, modest, and heterogeneous within task. Every headline was independently recomputed from per-item data by the Director; Section 3.4 records the scope of each re-derivation. Two of the five outcomes went against us, and they are the credibility of the other three; costs and limitations are reported alongside benefits.
Keywords: LoRA, parameter-efficient fine-tuning, weight space, gauge symmetry, model merging, pre-registration
Headline results
Five weight-only analyses, locked before the bank completed and run against it exactly once:
The artifact. A 480-adapter pre-registered bank — 2 families × 6 tasks × 40 seeds — trained on a single local GPU.
Task identity. Raw weights fail task identification (0.0792 / 0.1375 against chance 0.1667) while GL(r)-gauge-canonical features reach 1.0000. The obstacle is the parameterization, not the information.
Cross-family transfer. The pre-registered no-transfer prediction was refuted by its own triviality control: standardized transfer runs 0.7375–0.7833.
What carries the structure. Bridge swaps cost less than 0.001 nats; destroying the backbone costs +2.8 / +3.8.
Merge conflict. Weight-only merge-conflict prediction at group-aware AUC 0.995 / 0.962.
Honest shrinkage. A pilot correlation of r = 0.888 shrank to r = 0.300 at bank scale.
The paper was audited by a 117-finding adversarial ladder, and every number was independently re-derived by an isolated Director instance.
Artifacts
The paper source, the analysis code, the pre-registration cards, the dated rulings, and the full audit ledger are public in the rhombic repository. The extracted adapter features are published as a dataset; the bank itself is held locally with its manifest and trainer configuration recorded in the paper.
| Artifact | Where |
|---|---|
| Paper (PDF) | rhombic-asset1.pdf |
| Paper source (LaTeX) | paper/rhombic-asset1.tex |
| Repository | github.com/tasumermaf/rhombic |
| Adapter features | timotheospaul/rhombic-asset1-features |
| Audit ledger | 117 findings, dispositions on record |
| Full audit report | paper/audit/round-2/ |
Also from this program
Typed State Beats Prose (XR-001, July 2026) — a pre-registered, matched-budget measurement of numeric corruption under agent context compaction: prose re-encoding corrupts 36.4% of numeric facts where typed state blocks corrupt 9.4%, paired McNemar p = 3.524×10−21.