{"_id":"4allhuman","_rev":"4-41fe22fd361ae2c65f9d028e14b3e44c","name":"4allhuman","dist-tags":{"latest":"0.1.2"},"versions":{"0.1.0":{"name":"4allhuman","version":"0.1.0","_id":"4allhuman@0.1.0","maintainers":[{"name":"zeecares","email":"apachewang1990@gmail.com"}],"homepage":"https://github.com/zeecares/4allhuman#readme","bugs":{"url":"https://github.com/zeecares/4allhuman/issues"},"bin":{"4allhuman":"bin/4allhuman.ts"},"dist":{"shasum":"58f5a613c4c9f83d9a3268f03946b2d819d97e36","tarball":"https://registry.npmjs.org/4allhuman/-/4allhuman-0.1.0.tgz","fileCount":11,"integrity":"sha512-djSB6epS/oW4iIdfIER0SCq/INWIujhB4A1mAnOijYXiNlpkK9LPPlj2Ws0Ucv7MY9fyjk+y6+y7idZ+As2JLQ==","signatures":[{"sig":"MEYCIQC9iY6RmsZQ1QXsUXGtVOG+mZ/ZMUvLBjDfHwLAGE7QAAIhAO/dCgRolNSb2jYCEnIt62KEWBhGFrdoT3FqQf4jw63S","keyid":"SHA256:DhQ8wR5APBvFHLF/+Tc+AYvPOdTpcIDqOhxsBHRwC7U"}],"unpackedSize":90728},"type":"module","engines":{"node":">=22.18"},"scripts":{"cli":"node bin/4allhuman.ts","dev":"next dev","lint":"next lint","test":"node --test tests/*.test.ts","build":"next build","start":"next start","typecheck":"tsc --noEmit","update:crawlers":"node update-crawlers.mjs"},"_npmUser":{"name":"zeecares","email":"apachewang1990@gmail.com"},"deprecated":"Broken CLI bin (Node blocks TS type-stripping in node_modules) - use 0.1.1 instead","repository":{"url":"git+https://github.com/zeecares/4allhuman.git","type":"git"},"_npmVersion":"10.9.8","description":"Protect any creator's content from AI training in 60 seconds: scan, generate opt-outs (robots.txt, ai.txt, meta tags, EU Art.4(3) reservation), and sign content with provenance.","directories":{},"_nodeVersion":"22.23.2","dependencies":{"next":"^15.1.0","react":"^19.0.0","react-dom":"^19.0.0"},"_hasShrinkwrap":false,"devDependencies":{"typescript":"^5.7.0","@types/node":"^22.10.0","@types/react":"^19.0.0","@types/react-dom":"^19.0.0"},"_npmOperationalInternal":{"tmp":"tmp/4allhuman_0.1.0_1788979009466_0.3752509416029701","host":"s3://npm-registry-packages-npm-production"}},"0.1.1":{"name":"4allhuman","version":"0.1.1","_id":"4allhuman@0.1.1","maintainers":[{"name":"zeecares","email":"apachewang1990@gmail.com"}],"homepage":"https://github.com/zeecares/4allhuman#readme","bugs":{"url":"https://github.com/zeecares/4allhuman/issues"},"bin":{"4allhuman":"dist/bin/4allhuman.js"},"dist":{"shasum":"8bafe2a628f2983624d30703cf40f75695ad7450","tarball":"https://registry.npmjs.org/4allhuman/-/4allhuman-0.1.1.tgz","fileCount":11,"integrity":"sha512-fTc9Apn8ctjG/uLXjbtFvU+qO7ZLJDvTJB98iDfmTFX+UQQigTdnvbS/7vAb+9mhyyYMQi8GSddZ+f7QRNRXwQ==","signatures":[{"sig":"MEUCIQCUEqxYM/S18cdCSYBfmRYM/g/94Ga1MT2k1shrTQnTBQIgLxTaEl/UimP1+GVzrnhx+nPdbTwAfRydWVGsNsMuC40=","keyid":"SHA256:DhQ8wR5APBvFHLF/+Tc+AYvPOdTpcIDqOhxsBHRwC7U"},{"sig":"MEUCIQD9JcWyj5ddFBBrrJtZ8iZocYUxe/AUKeLIp1jTn5kgswIgCqMjp6uLcaC01iiGPyEEyzE0/oD6dIHIaQ6kyyBB1d8=","keyid":"SHA256:DhQ8wR5APBvFHLF/+Tc+AYvPOdTpcIDqOhxsBHRwC7U"}],"unpackedSize":87888},"type":"module","engines":{"node":">=22.18"},"scripts":{"cli":"node bin/4allhuman.ts","dev":"next dev","lint":"next lint","test":"node --test tests/*.test.ts","build":"next build","start":"next start","prepack":"tsc -p tsconfig.build.json","typecheck":"tsc --noEmit","update:crawlers":"node update-crawlers.mjs"},"_npmUser":{"name":"zeecares","email":"apachewang1990@gmail.com"},"repository":{"url":"git+https://github.com/zeecares/4allhuman.git","type":"git"},"_npmVersion":"10.9.8","description":"Protect any creator's content from AI training in 60 seconds: scan, generate opt-outs (robots.txt, ai.txt, meta tags, EU Art.4(3) reservation), and sign content with provenance.","directories":{},"_nodeVersion":"22.23.2","dependencies":{"next":"^15.1.0","react":"^19.0.0","react-dom":"^19.0.0"},"_hasShrinkwrap":false,"devDependencies":{"typescript":"^5.7.0","@types/node":"^22.10.0","@types/react":"^19.0.0","@types/react-dom":"^19.0.0"},"_npmOperationalInternal":{"tmp":"tmp/4allhuman_0.1.1_1788984188250_0.1840449010783085","host":"s3://npm-registry-packages-npm-production"}},"0.1.2":{"_id":"4allhuman@0.1.2","bin":{"4allhuman":"dist/bin/4allhuman.js"},"bugs":{"url":"https://github.com/zeecares/4allhuman/issues"},"dist":{"shasum":"1e51096fe3e00838dca0ddf8727852ea1221dc35","tarball":"https://registry.npmjs.org/4allhuman/-/4allhuman-0.1.2.tgz","fileCount":11,"integrity":"sha512-YEr98/li4txcpR1A4eY8aA1qGpDXYYxGDbBJo2BeK/YRiDqJFK6cde8WI1SmHQo3YtfsmykGQ0s3jXkO2MkJqA==","signatures":[{"sig":"MEYCIQDJr4Z9NFDlhmEwBDgjU5Viz19rQZsjtST2wp+CnR7WuwIhAMMpmiqIK6oK6nQMwyJCHjo3IIhjnpwvCues9J+6A7M/","keyid":"SHA256:DhQ8wR5APBvFHLF/+Tc+AYvPOdTpcIDqOhxsBHRwC7U"},{"keyid":"SHA256:DhQ8wR5APBvFHLF/+Tc+AYvPOdTpcIDqOhxsBHRwC7U","sig":"MEQCIG3u16mF2GmYCsvWFj30raxi59x5nZTrRB7nrHMIeyPPAiBa/SGP/jm8d97j5VvewLLI0X/wBKysdpVi6s+l/sezaA=="}],"attestations":{"url":"https://registry.npmjs.org/-/npm/v1/attestations/4allhuman@0.1.2","provenance":{"predicateType":"https://slsa.dev/provenance/v1"}},"unpackedSize":87887},"name":"4allhuman","type":"module","engines":{"node":">=22.18"},"gitHead":"8a8b364fa0ef92d64c8716ab7fbacbf4e597c18f","scripts":{"cli":"node bin/4allhuman.ts","dev":"next dev","lint":"next lint","test":"node --test tests/*.test.ts","build":"next build","start":"next start","prepack":"tsc -p tsconfig.build.json","typecheck":"tsc --noEmit","update:crawlers":"node update-crawlers.mjs"},"version":"0.1.2","_npmUser":{"name":"GitHub Actions","email":"npm-oidc-no-reply@github.com","trustedPublisher":{"id":"github","oidcConfigId":"oidc:60bf70a1-6284-4541-8638-7ca71f48cf11"}},"homepage":"https://github.com/zeecares/4allhuman#readme","repository":{"url":"git+https://github.com/zeecares/4allhuman.git","type":"git"},"_npmVersion":"11.19.0","description":"Protect any creator's content from AI training in 60 seconds: scan, generate opt-outs (robots.txt, ai.txt, meta tags, EU Art.4(3) reservation), and sign content with provenance.","directories":{},"maintainers":[{"name":"zeecares","email":"apachewang1990@gmail.com"}],"_nodeVersion":"24.20.0","dependencies":{"next":"^15.1.0","react":"^19.0.0","react-dom":"^19.0.0"},"_hasShrinkwrap":false,"devDependencies":{"typescript":"^5.7.0","@types/node":"^22.10.0","@types/react":"^19.0.0","@types/react-dom":"^19.0.0"},"_npmOperationalInternal":{"host":"s3://npm-registry-packages-npm-production","tmp":"tmp/4allhuman_0.1.2_1789022614464_0.9691907248802074"}}},"time":{"created":"2026-09-09T18:36:49.297Z","modified":"2026-09-10T06:43:34.894Z","0.1.0":"2026-09-09T18:36:49.618Z","0.1.1":"2026-09-09T20:03:08.345Z","0.1.2":"2026-09-10T06:43:34.558Z"},"bugs":{"url":"https://github.com/zeecares/4allhuman/issues"},"homepage":"https://github.com/zeecares/4allhuman#readme","repository":{"url":"git+https://github.com/zeecares/4allhuman.git","type":"git"},"description":"Protect any creator's content from AI training in 60 seconds: scan, generate opt-outs (robots.txt, ai.txt, meta tags, EU Art.4(3) reservation), and sign content with provenance.","maintainers":[{"name":"zeecares","email":"apachewang1990@gmail.com"}],"readme":"# 🛡️ Don't Train On Me\n\n**Protect any creator's content from AI training in 60 seconds.**\n\nPaste your domain → get a protection score → receive ready-to-deploy opt-out artifacts:\n\n1. **robots.txt** section blocking 170+ known AI crawlers (GPTBot, ClaudeBot, Google-Extended, CCBot, Bytespider…) — doubling as a machine-readable **EU DSM Directive Art. 4(3)** rights reservation\n2. **/ai.txt** — Spawning-spec machine-readable access policy\n3. **`<meta name=\"robots\" content=\"noai, noimageai\">`** tags for every page head\n4. **Legal notice** of reserved rights\n\n## Why it matters (the legal hook)\n\n- **EU DSM Directive Art. 4(3):** creators may expressly reserve their content from text-and-data-mining \"in an appropriate manner, such as machine-readable means.\"\n- **EU AI Act Art. 53(1)(c):** GPAI providers *must* identify and comply with those reservations — honoring this output is a regulatory obligation for them.\n- **US:** no opt-out statute; fair use is uncertain (USCO Part 3 report). Blocking crawlers + keeping content behind access control is currently the strongest lever.\n\nResearch notes:\n- [`research/protecting-human-content-from-ai-training.md`](./research/protecting-human-content-from-ai-training.md) — legal & technical landscape (EU DSM Art. 4(3), AI Act Art. 53(1)(c), USCO Part 3, Glaze/Nightshade, IETF aipref, C2PA)\n- [`research/enhancement-research.md`](./research/enhancement-research.md) — verified build opportunities: Common Crawl CDX corpus check, DE-COP memorization detection, RSL licensing standard, x402 payments\n\n## What this can and cannot guarantee\n\n**No tool can make already-public content untrainable.** Anyone can download what anyone\ncan read. Our artifacts are signals + legal reservations, not force fields:\n\n| Level | Mechanism | Strength |\n|---|---|---|\n| Never expose it | Login/paywall/app | ✅ True guarantee — unfetched bytes can't be trained on |\n| Legal reservation (EU) | robots.txt Art. 4(3) statement | ⚠️ Doesn't prevent copying; makes it *infringing* for EU-regulated GPAI providers |\n| Contract | No-training terms on private sharing | ⚠️ After-the-fact claim |\n| Public + signals only | robots.txt alone, no jurisdiction | ❌ Honor system |\n\nThis tool sits in rows 2–3: it measures exposure, perfects the strongest legal lever available,\nand verifies deployment. For a true guarantee, keep content behind access control.\n\nSee `research/protecting-human-content-from-ai-training.md` for sources.\n\n## Design principles (first-principles trust)\n\nThe product reduces an information asymmetry between creators and AI trainers. Trust\nfollows from three commitments baked into the UI:\n\n1. **Verifiability over assertion** — every scan result links to the raw evidence\n   (the live `robots.txt` we read), with a UTC timestamp. Users can re-do any check by hand.\n2. **Honest limits, stated up front** — the \"What this can / cannot do\" module says plainly:\n   signals bind compliant crawlers only; we cannot audit what a model was trained on; access\n   control is the only complete protection. Overclaiming is the fastest way to lose trust.\n3. **No black boxes** — fixed, published score weights; no accounts; nothing stored;\n   open-source methodology.\n\nVisual language follows Teenage Engineering's industrial design: flat panels, hard edges,\nwarm-grey chassis, signature orange accents, monospace labels, LED status indicators,\nnumbered modules — instruments, not marketing.\n\n## CLI\n\nRun the same audit from your terminal or CI - no config, zero dependencies:\n\n```bash\nnpx 4allhuman audit example.com\n```\n\n```\nnpx 4allhuman audit example.com [--json] [--fail-below <0-100>] [--timeout <ms>]\n\n  --json                Machine-readable report (all layers, verdicts, score).\n  --fail-below <0-100>  Exit 1 when the score is below this threshold (CI gate).\n  --timeout <ms>        Per-request timeout (default 8000).\n```\n\nThe report prints one plain-language verdict per layer (PROTECTED / PARTIAL /\nMISSING / NOTE) with the evidence behind it, then the total score out of 100.\n\nExit codes: `0` audit ran and meets `--fail-below` (if given) · `1` score below\n`--fail-below` · `2` usage error or the site could not be reached.\n\nRuns directly from TypeScript via Node's type stripping - requires Node >= 22.18.\nThe CLI and the web app share the exact same audit engine (`src/lib/scanner.ts` +\n`src/lib/layers.ts`), so a terminal score always matches a web score.\n\n## Run it\n\n```bash\nnpm install\nnpm run dev        # http://localhost:3000\nnpm run update:crawlers   # refresh the crawler blocklist from the ai.robots.txt community dataset\nnpm test               # RFC 9309 + multi-layer audit tests (node:test, zero deps)\nnpm run typecheck      # tsc --strict, no emit\n```\n\n## Architecture\n\n```\nbin/\n└── 4allhuman.ts          # CLI entry point (shebang, zero deps, Node >= 22.18)\nsrc/\n├── cli.ts                # CLI core: arg parsing, report rendering, exit codes\n├── app/\n│   ├── page.tsx              # scan UI + artifact viewer with copy buttons\n│   └── api/protect/route.ts  # orchestrator: scan → generate → score\n└── lib/\n    ├── crawlers.ts           # re-exports the crawler blocklist + legal text\n    │   ├── crawlers.generated.ts  # 170+ crawlers, generated from the community ai.robots.txt dataset\n    ├── robots9309.ts         # RFC 9309 matcher: group selection, longest-match, wildcards\n    ├── layers.ts             # multi-layer audit: X-Robots-Tag, TDMRep, aipref, llms.txt, meta\n    ├── scanner.ts            # live scan, per-crawler verdicts with rule-level reasons\n    └── generator.ts          # artifact generation + layered protection score\ntests/\n├── robots9309.test.ts        # spec-case tests: ties, wildcards, group precedence\n└── layers.test.ts            # multi-layer parsing + score tests\n```\n\n## Radar (module 09) — manual probe mode\n\nPaste an article (or its URL) → get generated probe questions → ask ChatGPT /\nPerplexity / Claude / Gemini yourself → paste their answers back → 8-gram plagiarism\nforensics run **locally in your browser**.\n\nNo API keys. No AI-company calls from our servers. Nothing stored. The evidence pack\n(.json download) seals the source text with SHA-256 and records every verdict with\nmatched spans.\n\nThresholds: ≥10% word-8-gram containment = **COPIED**, 2–9% = SUSPICIOUS, else CLEAN.\n8 consecutive shared words is the classical plagiarism-forensics threshold — chance\nco-occurrence is effectively zero.\n\n## The multi-layer audit\n\nOne report across every opt-out standard, each layer with its own plain-language\nverdict and evidence:\n\n| Layer | Weight | What we check |\n|---|---|---|\n| robots.txt (RFC 9309) | 40 | Live evaluation against 170+ known AI crawlers: most-specific group, longest-match, allow-wins ties, wildcards |\n| X-Robots-Tag header | 15 | `noai` / `noimageai` (global = full, per-bot or `none`/`noindex` = partial) |\n| noai meta tags | 15 | `<meta name=\"robots\" content=\"noai, noimageai\">` on the homepage |\n| TDMRep (W3C) | 10 | `tdm-reservation: 1` header or meta — the machine-readable EU DSM Art. 4(3) reservation |\n| ai.txt | 10 | `/ai.txt` published and actually denying training/dataset use (honor-based) |\n| aipref (IETF draft) | 5 | `ai-train=n`-style preference tokens (draft-ietf-aipref-vocab) |\n| Reachable | 5 | The site answered the scan, so every verdict is live evidence |\n| llms.txt | 0 | Informational only — it is the ALLOW-side counterpart (guides LLM use), not an opt-out |\n\nThe same layer builders score user-pasted artifacts in verify-your-fix, so the\nbefore/after score is computed identically to a live scan.\n\n## Modules\n\n| # | Module | What it does |\n|---|---|---|\n| 02 | Verdict | Multi-layer audit — one 0–100 score across every opt-out standard (weights below) |\n| 03 | Evidence | Live robots.txt check across 170+ AI crawlers (community ai.robots.txt dataset), with source links |\n| 04 | Common Crawl | Checks the 6 latest CC indexes for your domain — the open corpus most training sets build on |\n| 05–09 | The Fix | robots.txt (+ EU Art. 4(3) reservation + RSL `License:` line), verify-your-fix, ai.txt, meta tags, legal notice |\n| 10 | RSL License | `/license.xml` in RSL 1.0 schema — machine-readable licensing for the AI-first web |\n| 11 | Radar | Manual probe mode: generate questions → ask engines yourself → local 8-gram forensics; plus **DE-COP memorization quiz** (arXiv 2402.09910) scoring verbatim-recognition against the 25% chance line |\n| §§ | Honesty layer | What this can and cannot do |\n\n## Key references\n\n- Directive (EU) 2019/790 Arts. 3–4 (TDM opt-out) — [EUR-Lex](https://eur-lex.europa.eu/legal-content/EN/TXT/?uri=CELEX:32019L0790)\n- Regulation (EU) 2024/1689 (AI Act) Art. 53(1)(c)–(d) — [EUR-Lex](https://eur-lex.europa.eu/legal-content/EN/TXT/?uri=OJ:L_202401689)\n- U.S. Copyright Office, *Copyright and AI Part 3: Generative AI Training* (2025) — [copyright.gov/ai](https://www.copyright.gov/ai/)\n- DE-COP: detecting copyrighted content in training data — [arXiv 2402.09910](https://arxiv.org/abs/2402.09910)\n- Beyond Public Access in LLM Pre-Training Data — [arXiv 2505.00020](https://arxiv.org/abs/2505.00020)\n- Stealing Part of a Production Language Model (canary extraction) — [arXiv 2403.06634](https://arxiv.org/abs/2403.06634)\n- RSL 1.0 Really Simple Licensing — [rslstandard.org](https://rslstandard.org/) · [RSL Collective](https://rslcollective.org/)\n- IETF AI Preferences WG — [datatracker](https://datatracker.ietf.org/wg/aipref/) · C2PA — [c2pa.org](https://c2pa.org/)\n- Glaze / Nightshade — [UChicago SAND Lab](https://glaze.cs.uchicago.edu/) · Spawning DO NOT TRAIN — [spawning.ai](https://spawning.ai/)\n- Common Crawl corpus & index API — [commoncrawl.org](https://commoncrawl.org/)\n\n## Roadmap (see TASKS.md)\n\n- [x] Scanner, generators, scoring, UI\n- [ ] One-click deploy: open a PR to the site's repo / upload via FTP\n- [ ] C2PA \"Proof of Human\" content signing\n- [ ] Spawning DO NOT TRAIN registry submission\n\n\n","readmeFilename":"README.md"}