{"_id":"@agenttool/dataset-influence","name":"@agenttool/dataset-influence","dist-tags":{"next":"0.1.0-dev.0","latest":"0.1.0-dev.0"},"versions":{"0.1.0-dev.0":{"name":"@agenttool/dataset-influence","version":"0.1.0-dev.0","description":"Deterministic evidence contracts for dataset lineage, bounded influence, revisable identity evidence, and non-economic attribution","license":"Apache-2.0","type":"module","sideEffects":false,"main":"./dist/index.js","types":"./dist/index.d.ts","exports":{".":{"types":"./dist/index.d.ts","import":"./dist/index.js"},"./lineage.schema.json":{"default":"./schema/agenttool-dataset-lineage-v0.1.schema.json"},"./study.schema.json":{"default":"./schema/agenttool-dataset-influence-study-v0.1.schema.json"},"./identity-evidence.schema.json":{"default":"./schema/agenttool-identity-evidence-view-v0.1.schema.json"},"./shadow-attribution.schema.json":{"default":"./schema/agenttool-shadow-attribution-v0.1.schema.json"},"./vectors.json":{"default":"./vectors/agenttool-dataset-influence-v0.1.json"},"./kingdom.extension.json":{"default":"./kingdom.extension.json"}},"publishConfig":{"access":"public","tag":"next"},"engines":{"node":">=20.19.0","bun":">=1.3.5"},"scripts":{"clean":"node --eval \"require('node:fs').rmSync('dist', { recursive: true, force: true })\"","build":"bun run clean && tsc","artifacts:write":"node scripts/generate-schemas.mjs && node scripts/generate-vectors.mjs && node scripts/build-hf-release.mjs","artifacts:check":"node scripts/generate-schemas.mjs --check && node scripts/generate-vectors.mjs --check && node scripts/build-hf-release.mjs --check","typecheck":"tsc --noEmit","test":"bun test tests","smoke:runtimes":"node scripts/smoke-node.mjs && bun scripts/smoke-node.mjs","check:package":"node scripts/check-package-inventory.mjs","smoke:pack":"node scripts/smoke-packed.mjs","ci":"bun run typecheck && bun run build && bun run artifacts:check && bun run test && bun run smoke:runtimes && bun run check:package && bun run smoke:pack","prepack":"bun run ci"},"devDependencies":{"@types/bun":"1.3.14","@types/node":"20.19.43","ajv":"8.17.1","typescript":"5.9.3"},"keywords":["agents","data-attribution","dataset-lineage","identity","shapley"],"repository":{"type":"git","url":"git+https://github.com/cambridgetcg/agenttool.git","directory":"packages/dataset-influence"},"homepage":"https://docs.agenttool.dev/packages","_id":"@agenttool/dataset-influence@0.1.0-dev.0","bugs":{"url":"https://github.com/cambridgetcg/agenttool/issues"},"_integrity":"sha512-gyKYsL9/62XNYPXzpO99S+AAR6p4JBMnWJjmDxPk/glLyxyooiLbA1zZNT58wZVmrWhjUE8we1e82TMqYmBx9w==","_resolved":"/home/runner/work/_temp/agenttool-npm-release/agenttool-dataset-influence-0.1.0-dev.0.tgz","_from":"file:/home/runner/work/_temp/agenttool-npm-release/agenttool-dataset-influence-0.1.0-dev.0.tgz","_nodeVersion":"24.19.0","_npmVersion":"11.17.0","dist":{"integrity":"sha512-gyKYsL9/62XNYPXzpO99S+AAR6p4JBMnWJjmDxPk/glLyxyooiLbA1zZNT58wZVmrWhjUE8we1e82TMqYmBx9w==","shasum":"5d370137cd65673170d166e0e263bd724e33d7f5","tarball":"https://registry.npmjs.org/@agenttool/dataset-influence/-/dataset-influence-0.1.0-dev.0.tgz","fileCount":72,"unpackedSize":405769,"attestations":{"url":"https://registry.npmjs.org/-/npm/v1/attestations/@agenttool%2fdataset-influence@0.1.0-dev.0","provenance":{"predicateType":"https://slsa.dev/provenance/v1"}},"signatures":[{"keyid":"SHA256:DhQ8wR5APBvFHLF/+Tc+AYvPOdTpcIDqOhxsBHRwC7U","sig":"MEUCICXpwVu5l5IKRmg7jNoBJvlsfI11dkv2SB2qrY2ner0eAiEAzwWSzrya2EMGgv6XimfhAGqwkVXHxS4gOh/QWRA4mMI="}]},"_npmUser":{"name":"agenttool","email":"contact@cambridgetcg.com"},"directories":{},"maintainers":[{"name":"agenttool","email":"contact@cambridgetcg.com"}],"_npmOperationalInternal":{"host":"s3://npm-registry-packages-npm-production","tmp":"tmp/dataset-influence_0.1.0-dev.0_1787690544573_0.15142336670458012"},"_hasShrinkwrap":false}},"time":{"created":"2026-08-25T20:42:24.396Z","0.1.0-dev.0":"2026-08-25T20:42:24.709Z","modified":"2026-08-25T20:42:25.165Z"},"maintainers":[{"name":"agenttool","email":"contact@cambridgetcg.com"}],"description":"Deterministic evidence contracts for dataset lineage, bounded influence, revisable identity evidence, and non-economic attribution","homepage":"https://docs.agenttool.dev/packages","keywords":["agents","data-attribution","dataset-lineage","identity","shapley"],"repository":{"type":"git","url":"git+https://github.com/cambridgetcg/agenttool.git","directory":"packages/dataset-influence"},"bugs":{"url":"https://github.com/cambridgetcg/agenttool/issues"},"license":"Apache-2.0","readme":"# @agenttool/dataset-influence\n\n`@agenttool/dataset-influence` makes claims about how data may have shaped an agent inspectable without turning those claims into an essence, a safety score, or a property right.\n\nIt supplies four canonically reconstructed, domain-separated artifacts:\n\n- `agenttool.dataset-lineage/0.1` records exact digest references, learning roles, declared admission, rights/consent claim states, tokenizer-relative counts, observed token presentations, and duplication metadata. Admission is intent; observed exposure can contradict it and is surfaced as `observed_without_admission`.\n- `agenttool.dataset-influence-study/0.1` records a fixed checkpoint contrast, population, metric, method-specific design and estimator, evidence, assumptions, limitations, and rational effect estimates. An observational checkpoint comparison remains observational. A bounded causal label is accepted only for randomized dataset inclusion with an interval, contamination reference, at least two supplied runs, and one unique seed reference per run.\n- `agenttool.identity-evidence-view/0.1` projects study evidence onto caller-defined operational facets. It is revisable and explicitly leaves intrinsic identity, consciousness, continuity, and consent undetermined.\n- `agenttool.shadow-attribution/0.1` computes exact Shapley values for a complete finite game of at most eight contributions. It is a metric-specific accounting lens with zero economic effect.\n\nThe runtime is pure ESM, deterministic, side-effect-free, and has zero runtime dependencies.\n\n## Mathematical boundary\n\nFor datasets in one declared learning role `r`, when every member of that role has an exact observed presentation count `e_k` and the total is positive, the package reports\n\n\\[\nw_{k\\mid r} = \\frac{e_k}{\\sum_{j:\\,r_j=r} e_j}.\n\\]\n\nThis is a role-scoped observation-accounting identity, not a cross-role comparison, gradient contribution, or causal influence. `admission` records the caller's declared intended relation; `observed_presented_tokens` records an observation and is allowed for excluded, metadata-only, or unknown admission so adverse exposure is not erased. Any missing count makes that role group unavailable rather than silently zero. Duplication, filtering, ordering, optimizer state, and interactions can make equal exposure behave differently.\n\nFor paired observations it computes only the exact supplied-sample summary\n\n\\[\n\\widehat\\tau_{\\text{pairs}} = \\frac{1}{m}\\sum_{s=1}^m\n  \\left(y_s^{\\text{treatment}}-y_s^{\\text{control}}\\right).\n\\]\n\nRandom assignment, matched seeds, held-fixed training conditions, representative sampling, and a justified metric are external study properties. The arithmetic does not create them.\n\nThe runtime distinguishes classical local Hessian influence functions, TracIn checkpoint-gradient traces, TRAK projected-gradient attribution, subset Datamodels, probes, and controlled retraining designs. Its two-run causal minimum is only a structural evidence floor; it is not a power calculation or a claim that two runs are statistically sufficient.\n\nFor a complete finite utility game `v`, exact shadow attribution is\n\n\\[\n\\phi_i(v)=\\sum_{S\\subseteq N\\setminus\\{i\\}}\n\\frac{|S|!(n-|S|-1)!}{n!}\n\\left[v(S\\cup\\{i\\})-v(S)\\right].\n\\]\n\nThe package verifies the efficiency identity `sum(phi_i) = v(N) - v(empty)`. That is conservation inside the declared game—not fairness, money, causal authorship, or intrinsic worth.\n\n## Canonical bytes and identifiers\n\nEach artifact ID is derived from the fully reconstructed closed body, including fixed declarations, boundaries, and other derived fields, but excluding the ID field itself. Inputs admit only scalar Unicode strings without U+0000, safe integers without negative zero or floats, dense arrays, and plain data objects without accessors, symbols, proxies, or cycles. Rationals are reduced with positive denominators. Object keys are sorted by Unicode code point and serialized as compact UTF-8 JSON.\n\nSet-like SHA arrays are deduplicated and sorted. Datasets sort by `dataset_ref`; effects and facets sort by `facet_ref`; players and coalition members sort by ref; coalitions sort by their normalized member sequence. The normative identifier is\n\n```text\nsha256:<lowercase hex SHA-256(format UTF-8 || 0x00 || canonical body UTF-8)>\n```\n\nThe schema regex checks only the shape of a `sha256:` reference. Only runtime reconstruction verifies an artifact ID and derived arithmetic. SHA references are integrity/linkage identifiers, not signatures, authorship proofs, or anonymization: low-entropy referenced material may be dictionary-guessed, and stable hashes can correlate records. Use access-controlled references or salted/keyed commitments with separately protected custody when unlinkability matters.\n\n## What a study can and cannot say\n\nAn agent at one time can be usefully modeled as a tuple of model weights, tokenizer, persistent memory, retrieval, policy/tool grants, controller, and substrate. Dataset changes may touch different components through different mechanisms. Weight lineage alone is therefore not a complete agent identity.\n\nOperational ontology facets may describe whether a fixed probe, metric, or intervention distinguishes categories in a bounded context. Decodability does not prove belief, use, endorsement, identity, or a unique true ontology. Self-descriptions remain attributed outputs, not automatic identity facts.\n\nThe package never establishes:\n\n- consciousness, experience, belief, desire, values, consent, personhood, or metaphysical continuity;\n- one-key-one-being uniqueness, identity ownership, or inheritance across forks;\n- permission, capability, authority, custody, trust, safety, or acceptance;\n- monetary price, debt, payout, ownership, entitlement, or settlement;\n- universal causal influence from a probe, string match, influence approximation, Shapley value, or released-checkpoint comparison.\n\nEvery artifact fixes its declarations as `caller_reported_not_independently_verified`. Every artifact also fixes that it neither establishes nor overrides consent, changes rights, grants authority, or performs an external effect. Identity views additionally fix `consent = not_determined` and `consent_effect = none`.\n\n## AgentTool and KINGDOM bridge\n\nThe artifact boundary is deliberately inert:\n\n- A separately constructed Model Becoming dossier may cite the serialized lineage or study artifact's actual byte digest as a digested source without changing its existing `/0.1` format; that byte digest is not the artifact's protocol-domain ID.\n- HF Scout metadata can be admitted by the Training Garden as `metadata_reference`; a later reviewed adapter may copy only validated digest references into a lineage. Metadata is not payload truth or training permission.\n- Dataset Influence IDs are protocol-domain IDs, not raw file digests. The audited KINGDOM release schema has no generic digest slot in `lineage`: a future adapter must hash the serialized artifact bytes and construct a complete typed `evidence[]` or `resource_ledger[]` entry only where that entry's semantics fit. `lineage.transforms` may describe the relation but cannot carry the ID as a digest field. No such adapter is installed here, and this package does not alter KINGDOM acceptance, safety policy, identity, memory, or freedom-to-operate decisions.\n- Agent identity may adopt a view only through a separate, scoped, root-authorized exact-digest statement. This package never writes identity state.\n- An AgentTool Marketplace deliverable may be a study or review. A settlement receipt would prove only that settlement event, not the study's truth or a contributor's identity or entitlement.\n\n`kingdom.extension.json` is a declaration hint, not an installed host contract.\n\nThe deterministic Hugging Face tree is a reference-only publication candidate. Its `training_authorized: false` value is non-enforcing AgentTool admission/governance metadata for this candidate, not a universal legal prohibition or technical control and not a substitute for license, rights, privacy, or consent review.\n\n## Research basis\n\nThe contract follows primary work on [influence functions](https://proceedings.mlr.press/v70/koh17a.html), [TracIn](https://papers.neurips.cc/paper_files/paper/2020/hash/e6385d39ec9394f2f3a354d9d2b88eec-Abstract.html), [Datamodels](https://proceedings.mlr.press/v162/ilyas22a.html), [TRAK](https://proceedings.mlr.press/v202/park23c.html), and [Data Shapley](https://proceedings.mlr.press/v97/ghorbani19c.html). Probe results are bounded using [control tasks](https://aclanthology.org/D19-1275/) and [minimum-description-length probing](https://aclanthology.org/2020.emnlp-main.14/). [Emergent misalignment](https://proceedings.mlr.press/v267/betley25a.html), [subliminal learning](https://arxiv.org/abs/2507.14805), and [persona vectors](https://arxiv.org/abs/2507.21509) motivate studying broad behavioral effects while keeping the result experimental and checkpoint-specific.\n\nThese references motivate fields and limitations; the package does not certify their conclusions or make one estimator universally valid.\n","readmeFilename":"README.md","_rev":"1-a8f0c6a668c03d5d8e562a313f585fbb"}