Governed clinical variable mapping with browser-local reasoning
Mappings proposed, not applied — 426 nodes, 332 edges, 0 of 84 approved. A mapping waits for a reviewer.
Mapping study variables to analysis variables is slow, manual and repeated per study. A model can propose mappings, but a proposal accepted because the model sounded confident is not a mapping — it is a guess with provenance attached.
Built on the public CDISC pilot: embedding retrieval over a 415-variable candidate pool against 110 gold pairs, with identical cached embeddings across every run so comparisons mean something. Candidate scoring runs in-page over WebGPU, so the data does not leave the browser. A mapping is always a proposal a reviewer accepts, never something applied automatically.
On the derivation split — the hard half — top-1 accuracy is 0.400 for a frontier model against 0.220 lexical and 0.140 on-premises. Across all 110 pairs the spread narrows considerably, which is why the split is the number worth quoting.
Four-page paper under review at IEEE BIBM 2026.
No public URL — paper under review.