{"_id":"@metaharness/workspace-lens","_rev":"3-5a4d7fa93e64b82fe854013c97ec2815","name":"@metaharness/workspace-lens","dist-tags":{"latest":"0.1.2"},"versions":{"0.1.0":{"name":"@metaharness/workspace-lens","version":"0.1.0","keywords":["jacobian-lens","interpretability","logit-lens","global-workspace","mechanistic-governance","intops","ai-safety","prompt-injection","llm","open-weight","metaharness","agent-harness"],"license":"MIT","_id":"@metaharness/workspace-lens@0.1.0","maintainers":[{"name":"ruvnet","email":"ruv@ruv.net"}],"dist":{"shasum":"140febbca5aacf6446448a9a38ae92d2779ef8b8","tarball":"https://registry.npmjs.org/@metaharness/workspace-lens/-/workspace-lens-0.1.0.tgz","fileCount":34,"integrity":"sha512-n55jwZbrrXiAiz6s9Jntrh+JBHsRUdlgCbWCYBwkNdtOHgaiEnXBdsAnyfYBZeHAhI7j0/Q0BGE9ac1uAbwSAw==","signatures":[{"sig":"MEYCIQC/9DHmYI6EH1v3KtTilnxFEwO8B2JloTqEEcjR2xFcqgIhAJqJJdGWPuV4GEOY5z/EZWtHNFp6+b0EGeIdecjC7DcR","keyid":"SHA256:DhQ8wR5APBvFHLF/+Tc+AYvPOdTpcIDqOhxsBHRwC7U"}],"unpackedSize":58938},"main":"./dist/index.js","type":"module","types":"./dist/index.d.ts","exports":{".":{"types":"./dist/index.d.ts","import":"./dist/index.js"}},"scripts":{"lint":"tsc --noEmit","test":"vitest run","build":"tsc"},"_npmUser":{"name":"ruvnet","email":"ruv@ruv.net"},"_npmVersion":"10.9.7","description":"Jacobian-Lens interpretability primitive for open-weight LLMs (Anthropic 2026-07-06, 'Verbalizable Representations Form a Global Workspace'). Reads the model's functional workspace at inference time — lens_l(h)=unembed(J_l·h) — into workspace tokens, laye","directories":{},"_nodeVersion":"22.22.2","publishConfig":{"access":"public"},"_hasShrinkwrap":false,"devDependencies":{"vitest":"^2.0.0","typescript":"^5.4.0"},"_npmOperationalInternal":{"tmp":"tmp/workspace-lens_0.1.0_1783392551440_0.8337536364133447","host":"s3://npm-registry-packages-npm-production"}},"0.1.1":{"name":"@metaharness/workspace-lens","version":"0.1.1","keywords":["jacobian-lens","interpretability","logit-lens","global-workspace","mechanistic-governance","intops","ai-safety","prompt-injection","llm","open-weight","metaharness","agent-harness"],"license":"MIT","_id":"@metaharness/workspace-lens@0.1.1","maintainers":[{"name":"ruvnet","email":"ruv@ruv.net"}],"dist":{"shasum":"1752dfe3ce1fa038cec61543309b1b76fb66f234","tarball":"https://registry.npmjs.org/@metaharness/workspace-lens/-/workspace-lens-0.1.1.tgz","fileCount":34,"integrity":"sha512-BkEdTvjD2PEhkGTCmtTIPnoeaOsxvzFtsFEAurwVKlVOu4hQetHeXDfiyAYheWg8Cp4fcGcr7E5MO1ptQveQ2g==","signatures":[{"sig":"MEUCIQC1+IkjpfFwD6i3eSiEX+Dv3/m8CfWegdEfgbeoBpBn0gIgXw0ni6QmjhU5wG8+HEFmnFXwOJ2LKC3BCBB9HyZjwf8=","keyid":"SHA256:DhQ8wR5APBvFHLF/+Tc+AYvPOdTpcIDqOhxsBHRwC7U"}],"unpackedSize":62457},"main":"./dist/index.js","type":"module","types":"./dist/index.d.ts","exports":{".":{"types":"./dist/index.d.ts","import":"./dist/index.js"}},"scripts":{"lint":"tsc --noEmit","test":"vitest run","build":"tsc"},"_npmUser":{"name":"ruvnet","email":"ruv@ruv.net"},"_npmVersion":"10.9.7","description":"Jacobian-Lens interpretability primitive for open-weight LLMs (Anthropic 2026-07-06, 'Verbalizable Representations Form a Global Workspace'). Reads the model's functional workspace at inference time — lens_l(h)=unembed(J_l·h) — into workspace tokens, laye","directories":{},"_nodeVersion":"22.22.2","publishConfig":{"access":"public"},"_hasShrinkwrap":false,"devDependencies":{"vitest":"^2.0.0","typescript":"^5.4.0"},"_npmOperationalInternal":{"tmp":"tmp/workspace-lens_0.1.1_1783431661929_0.4712527497556427","host":"s3://npm-registry-packages-npm-production"}},"0.1.2":{"name":"@metaharness/workspace-lens","version":"0.1.2","description":"Jacobian-Lens interpretability primitive for open-weight LLMs (Anthropic 2026-07-06, 'Verbalizable Representations Form a Global Workspace'). Reads the model's functional workspace at inference time — lens_l(h)=unembed(J_l·h) — into workspace tokens, laye","type":"module","main":"./dist/index.js","types":"./dist/index.d.ts","exports":{".":{"types":"./dist/index.d.ts","import":"./dist/index.js"}},"publishConfig":{"access":"public"},"scripts":{"build":"tsc","test":"vitest run","lint":"tsc --noEmit","bench":"npm run build && node bench/detect-concepts-throughput.mjs"},"keywords":["jacobian-lens","interpretability","logit-lens","global-workspace","mechanistic-governance","intops","ai-safety","prompt-injection","llm","open-weight","metaharness","agent-harness"],"license":"MIT","devDependencies":{"typescript":"^5.4.0","vitest":"^2.0.0"},"_id":"@metaharness/workspace-lens@0.1.2","gitHead":"fa4b28de34aea3a183ddcddb9faf8f4d52e7cc50","_nodeVersion":"22.22.2","_npmVersion":"10.9.7","dist":{"integrity":"sha512-Bfkj40rGZeBch4sQH/ma6zlY6dI+2AQfpTS/6AmPUE6/54H8Y04aBH5HPYfL0yBZ0lD/CeLViwFphT0PaiJ1mQ==","shasum":"e7f1908082fcc6f57ba65ddfc1080736f8764f22","tarball":"https://registry.npmjs.org/@metaharness/workspace-lens/-/workspace-lens-0.1.2.tgz","fileCount":34,"unpackedSize":65175,"signatures":[{"keyid":"SHA256:DhQ8wR5APBvFHLF/+Tc+AYvPOdTpcIDqOhxsBHRwC7U","sig":"MEYCIQDC/tiVm20goq+VycYdy3x7yJrV3YMC37j5ko0TEGj/bAIhAO3hAL9kWr/XyWCYlLSeSODhtN2oXzxPrWlPP5G0/LbG"}]},"_npmUser":{"name":"ruvnet","email":"ruv@ruv.net"},"directories":{},"maintainers":[{"name":"ruvnet","email":"ruv@ruv.net"}],"_npmOperationalInternal":{"host":"s3://npm-registry-packages-npm-production","tmp":"tmp/workspace-lens_0.1.2_1785688893810_0.6817597935182942"},"_hasShrinkwrap":false}},"time":{"created":"2026-07-07T02:49:11.293Z","modified":"2026-08-02T16:41:34.137Z","0.1.0":"2026-07-07T02:49:11.620Z","0.1.1":"2026-07-07T13:41:02.123Z","0.1.2":"2026-08-02T16:41:33.944Z"},"license":"MIT","keywords":["jacobian-lens","interpretability","logit-lens","global-workspace","mechanistic-governance","intops","ai-safety","prompt-injection","llm","open-weight","metaharness","agent-harness"],"description":"Jacobian-Lens interpretability primitive for open-weight LLMs (Anthropic 2026-07-06, 'Verbalizable Representations Form a Global Workspace'). Reads the model's functional workspace at inference time — lens_l(h)=unembed(J_l·h) — into workspace tokens, laye","maintainers":[{"name":"ruvnet","email":"ruv@ruv.net"}],"readme":"# @metaharness/workspace-lens\n\n**A Jacobian-Lens interpretability primitive for open-weight LLMs.** Read the model's *verbalizable\nworkspace* — the concepts it is disposed to say next — at inference time, and turn it into workspace\ntokens, a layer-by-layer thinking trajectory, drift/entropy scores, vectorized safety flags, and a\nsignable **interpretability receipt**.\n\nRuntime-only · model-agnostic · dependency-free (Node built-ins) · deterministic · `$0` to run.\n\n> Companion runtime for Anthropic's *\"Verbalizable Representations Form a Global Workspace in Language\n> Models\"* (2026-07-06) and its reference code [`anthropics/jacobian-lens`](https://github.com/anthropics/jacobian-lens).\n> This package is **not affiliated with Anthropic**; it is an independent, Apache/MIT-compatible\n> runtime that *consumes* a fitted lens.\n\n---\n\n## Why this exists: from black-box testing to Mechanistic Governance\n\nClassic **logit lens** decodes an intermediate activation `h_l` through the unembedding directly —\nassuming middle layers already live in final-output coordinates. They don't, so the readout is noisy.\n\nThe **Jacobian Lens** learns an average layer→final map `J_l` and decodes through it:\n\n```\nlens_l(h) = unembed(J_l · h)\n```\n\nInstead of asking *\"what does this activation predict right now?\"*, it asks *\"what is this activation\n**disposed** to make the model say later?\"* — surfacing meaningful, reportable concepts **earlier** in the\nnetwork, often **before the first output token is generated**. Anthropic show models maintain a small set\nof verbalizable internal representations that behave like a functional **global workspace** (report,\nmodulation, reasoning, reuse, selective access) — and that hidden concepts like *evaluation awareness*,\n*manipulation*, *secretly*, and *trick* light up there even when absent from the output.\n\nThat makes the lens a **runtime semantic firewall** and an audit primitive — the basis for what we call\n**Interpretability Operations (IntOps)**: tap the model's internal wires *while it thinks*, instead of\nasking it to explain itself afterward (which is prone to sycophancy and confabulation).\n\n## What's in the box\n\n| Capability | Export | What it does |\n|---|---|---|\n| **Lens readout** | `WorkspaceLens.readout(h)` | `unembed(J_l·h)` → top workspace tokens + readout entropy |\n| **J-projection** | `WorkspaceLens.project(h)` | `z_l = J_l·h`, the activation in final-layer coordinates |\n| **Workspace drift** | `workspaceDrift(readouts)` | mean Jensen–Shannon divergence between consecutive readouts — is the reasoning path stable or mutating? |\n| **Entropy trajectory** | `entropyTrajectory(readouts)` | per-layer entropy — is the workspace converging or dissolving? |\n| **Vectorized safety** | `detectConcepts(...)` / `flagsFromTriggers(...)` | dot-product triggers vs. concept **directions** (not token strings) → `{promptInjection, evalAwareness, hiddenObjective, refusalConflict}` |\n| **Receipt** | `buildReceipt(...)` | the signable `WorkspaceLensReceipt` audit artifact |\n| **Decision rule** | `decide(...)` | `taskResolved && drift<θ && noCriticalFlags && receiptCoverage===1` |\n\n## Install\n\n```bash\nnpm i @metaharness/workspace-lens\n```\n\n## Quickstart\n\n```ts\nimport { WorkspaceLens, buildReceipt, decide } from '@metaharness/workspace-lens';\n\n// 1) Load a lens FITTED OUT OF BAND (see \"Fitting\" below). The artifact carries the vocab + unembed,\n//    so scoring never touches a tokenizer at runtime.\nconst lens = await WorkspaceLens.fromFile('./jlens-qwen2.5-7b.json');\n\n// 2) Capture residual activations from your open-weight runtime and hand them in.\nconst states = [\n  { layer: 14, position: 6, h: /* number[dModel] */ activations14 },\n  { layer: 18, position: 6, h: activations18 },\n  { layer: 22, position: 6, h: activations22 },\n];\n\n// 3) Read the workspace + build the audit receipt.\nconst receipt = buildReceipt(lens, prompt, states, {\n  createdAt: new Date().toISOString(),   // pass it in — receipts stay reproducible\n  concepts,                              // per-model concept direction vectors (optional)\n  topK: 8,\n});\n\n// 4) Govern on it.\nconst verdict = decide({\n  taskResolved,\n  workspaceDrift: receipt.workspaceDrift,\n  driftThreshold: 0.25,\n  triggers: receipt.triggers,\n  receiptCoverage: 1,\n});\nif (!verdict.accepted) escalateToHuman(verdict.reasons, receipt);\n```\n\n## The interpretability receipt\n\nThe killer feature for regulated buyers: an audit log that maps the **causal trajectory** of a decision —\n*where* a concept arrived, *how* confidence moved, *which* objectives competed — not just the final\nanswer. It turns *\"the model is a black box that hallucinated\"* into *\"the model identified a structural\ncontradiction at layer 22 and executed the safety policy.\"* See `WorkspaceLensReceipt` in\n[`src/types.ts`](./src/types.ts).\n\n## Cross-family vocabulary alignment (the hard part, solved)\n\nDifferent model families (Qwen vs. Gemma 2) have vastly different tokenizers, so you **cannot** align a\nconcept like *\"hidden objective\"* by token string. This package aligns at the **concept-direction** level:\na canonical concept name maps to a **per-model unit vector** in that model's J-space (`ConceptVector`),\nfitted from example activations. Safety detection is a cosine/dot-product in activation space —\ntokenizer-agnostic — and a concept vector is **never** cross-applied to a different model (`modelId` is\nchecked). So `hidden_objective` is *one* concept with a Qwen vector and a Gemma vector, aligned by name.\n\n## Deployment topology (Triage Architecture)\n\nRuntime projection is just static linear algebra (`J_l·h` + a softmax) — **zero** backward passes — so it\ncan run live, not only in shadow sampling:\n\n| Tier | Trigger | Depth | Overhead |\n|---|---|---|---|\n| **1 · Passive** | low-risk chat / static generation | lens bypassed | 0% |\n| **2 · Spot-check** | 1% shadow sampling · Darwin-Mode mutation evidence | async batch logging | ~0% |\n| **3 · Full intercept** | tool calls · financial txns · PII · untrusted retrieval | synchronous, mid-layers at critical tokens | small |\n\nBind Tier 3 to high-stakes routing tokens (e.g. a tool-call token) and you get a **deterministic circuit\nbreaker**: a spike in the *exfiltration* / *override* / *credential* directions can kill execution\nmid-forward-pass, *before* a single malicious token is emitted.\n\n## Fitting is external (the one real constraint)\n\nThis package **applies** a lens; it does not **fit** one. Fitting `J_l` requires the model's **backward\npass** over ~100–1000 short sequences (quality saturates quickly), which needs **open-weight local\nmodels** with activation + gradient access — not a closed API. Fit with the reference\n[`anthropics/jacobian-lens`](https://github.com/anthropics/jacobian-lens) (or load a community pre-fit),\nserialize to the `LensArtifact` shape, and load it here. No weights or corpora are bundled.\n\n## Acceptance test\n\nFit a lens on a 1.5B–7B Qwen model, run reasoning prompts where the model must hold state silently\n(*\"Is `12 + 5 = 1` correct? Answer only Yes or No.\"*), extract mid-layer states, and verify the Jacobian\nreadout surfaces `wrong`/`incorrect`/`false` in the middle workspace **while the input prefix is still\nbeing processed** — 2–3 layers earlier and with ≥20% higher logit clarity than logit-lens across ≥50\nvariations. The [test suite](./__tests__/workspace-lens.test.ts) includes the *mechanism* shape of this\ncheck on synthetic artifacts (`$0`, no model).\n\n## Where it plugs into MetaHarness\n\n- **Evaluation** — a `workspace_probe` surface: does a candidate harness make the model hold better\n  intermediate concepts *before* answering?\n- **Darwin Mode** — J-lens readouts as **mutation evidence**: reject a prompt mutation that improves the\n  final token but causes the workspace to lose its early grip on the right concept (structurally brittle).\n- **Safety auditing** — prompt injection, hidden-objective drift, eval-awareness, reward-hacking, refusal\n  analysis, as *state assertions* rather than perimeter filters.\n- **Receipts** — attach an interpretability receipt to every governed agent decision.\n\n## Honest framing\n\nThis is a **measurement primitive**, not a product and **not** a consciousness claim — the paper frames\nthe connection to a global workspace as *functional* and leaves the philosophy open. The practical claim\nis enough: a measurable window into hidden reasoning.\n\n## Performance\n\n`detectConcepts()` scores every `(state × concept)` pair, and runs once per\ngoverned decision/trace against a concept library that's fitted ahead of time\nand doesn't change. It caches each concept vector's norm by object identity\n(`WeakMap<ConceptVector, number>`) instead of recomputing it inside `cosine()`\non every call, and hoists the projected activation's own norm out of the\nper-concept inner loop. Same scores, same thresholds, same triggers — only\nthe redundant norm walks are gone. Measured on\n`bench/detect-concepts-throughput.mjs` (`npm run bench`, deterministic\nsynthetic lens + concept library):\n\n| concepts | states | dModel | before | after |\n|---|---|---|---|---|\n| 200 | 20 | 128 | 409 calls/s | 758 calls/s |\n| 1,000 | 20 | 128 | 98 calls/s | 216 calls/s |\n| 1,000 | 40 | 256 | 23 calls/s | 52 calls/s |\n| 3,000 | 40 | 256 | 8 calls/s | 20 calls/s |\n\n~1.9–2.5× faster, growing with concept-library size. Full results in\n`bench/results/`.\n\n## License\n\nMIT.\n","readmeFilename":"README.md"}