{"_id":"@daydreamsai/synthetic","_rev":"2-976b93042f7b97ff6da88896f37536f4","name":"@daydreamsai/synthetic","dist-tags":{"latest":"0.3.9-alpha.1"},"versions":{"0.3.9-alpha.0":{"name":"@daydreamsai/synthetic","version":"0.3.9-alpha.0","keywords":["ai","agents","synthetic-data","training","machine-learning","fine-tuning"],"author":{"name":"Daydreams"},"license":"MIT","_id":"@daydreamsai/synthetic@0.3.9-alpha.0","maintainers":[{"name":"realm_lord","email":"realm.lord.eth@gmail.com"},{"name":"bmgalego","email":"bmgalego@gmail.com"}],"dist":{"shasum":"05f6a13b4b0c1944fb4f36a153e41a954abdd259","tarball":"https://registry.npmjs.org/@daydreamsai/synthetic/-/synthetic-0.3.9-alpha.0.tgz","fileCount":8,"integrity":"sha512-hdn5TN/rPAaQhgs75N6wV5AZSSFO2ji2GJAsEpDXXiiRJCLtH3QpdWPT3cjP0ufCoWTVHBb2q6KHztP4b096sg==","signatures":[{"sig":"MEYCIQCp4FCwZUH+P0JyyyrDvV6Kr9g+Gjv118JJDg0WIokQaQIhAI8x/31ZFF8C+Nfkn19NRmBAxWe72wauEqZ17HqJtlvs","keyid":"SHA256:DhQ8wR5APBvFHLF/+Tc+AYvPOdTpcIDqOhxsBHRwC7U"}],"unpackedSize":345857},"main":"dist/index.js","_from":"file:daydreamsai-synthetic-0.3.9-alpha.0.tgz","types":"dist/index.d.ts","scripts":{"dev":"tsup --watch","test":"vitest","build":"tsup","test:watch":"vitest --watch","test:coverage":"vitest --coverage"},"_npmUser":{"name":"realm_lord","email":"realm.lord.eth@gmail.com"},"_resolved":"/tmp/4d0ba9d233b57d29263ed8b0ebb7a26a/daydreamsai-synthetic-0.3.9-alpha.0.tgz","_integrity":"sha512-hdn5TN/rPAaQhgs75N6wV5AZSSFO2ji2GJAsEpDXXiiRJCLtH3QpdWPT3cjP0ufCoWTVHBb2q6KHztP4b096sg==","_npmVersion":"10.8.2","description":"Synthetic data generation for AI agent training","directories":{},"_nodeVersion":"20.19.4","dependencies":{"zod":"3.25.23","@daydreamsai/core":"0.3.9-alpha.0"},"_hasShrinkwrap":false,"devDependencies":{"tsup":"8.3.6","vitest":"3.0.5","typescript":"5.8.3","@vitest/coverage-v8":"^3.2.0"},"peerDependencies":{"@daydreamsai/core":"0.3.9-alpha.0"},"_npmOperationalInternal":{"tmp":"tmp/synthetic_0.3.9-alpha.0_1754614567969_0.8501171048799101","host":"s3://npm-registry-packages-npm-production"}},"0.3.9-alpha.1":{"name":"@daydreamsai/synthetic","version":"0.3.9-alpha.1","description":"Synthetic data generation for AI agent training","main":"dist/index.js","types":"dist/index.d.ts","keywords":["ai","agents","synthetic-data","training","machine-learning","fine-tuning"],"author":{"name":"Daydreams"},"license":"MIT","dependencies":{"zod":"4.0.16","@daydreamsai/core":"0.3.9-alpha.1"},"devDependencies":{"@vitest/coverage-v8":"^3.2.0","tsup":"8.3.6","typescript":"5.8.3","vitest":"3.0.5"},"peerDependencies":{"@daydreamsai/core":"0.3.9-alpha.1"},"scripts":{"build":"tsup","dev":"tsup --watch","test":"vitest","test:watch":"vitest --watch","test:coverage":"vitest --coverage"},"_id":"@daydreamsai/synthetic@0.3.9-alpha.1","_integrity":"sha512-Wes7pyh0gO653E2NnQvAa0NZw3cud3QUxEzO5yLdFtl0HCxtkEZD1Ts4WEhj4TUy76dTW6kRpbNMEP6uTylRPw==","_resolved":"/tmp/b29c4f7a9bb5ebc3c65f72ab489fc58b/daydreamsai-synthetic-0.3.9-alpha.1.tgz","_from":"file:daydreamsai-synthetic-0.3.9-alpha.1.tgz","_nodeVersion":"20.19.4","_npmVersion":"10.8.2","dist":{"integrity":"sha512-Wes7pyh0gO653E2NnQvAa0NZw3cud3QUxEzO5yLdFtl0HCxtkEZD1Ts4WEhj4TUy76dTW6kRpbNMEP6uTylRPw==","shasum":"7dca62bef9af492a7029a98201ea9e8a4045add9","tarball":"https://registry.npmjs.org/@daydreamsai/synthetic/-/synthetic-0.3.9-alpha.1.tgz","fileCount":8,"unpackedSize":345826,"signatures":[{"keyid":"SHA256:DhQ8wR5APBvFHLF/+Tc+AYvPOdTpcIDqOhxsBHRwC7U","sig":"MEUCIQC5kyFKAUBKv9t4LHYlBYVcjji/blqpGQDCzI1+NzMhXAIgR0LI8pe4ksh43dAyw1ldIP7me4tQT/+EXCJZJNEgH9Q="}]},"_npmUser":{"name":"realm_lord","email":"realm.lord.eth@gmail.com"},"directories":{},"maintainers":[{"name":"realm_lord","email":"realm.lord.eth@gmail.com"},{"name":"bmgalego","email":"bmgalego@gmail.com"}],"_npmOperationalInternal":{"host":"s3://npm-registry-packages-npm-production","tmp":"tmp/synthetic_0.3.9-alpha.1_1754717902332_0.9469048186534923"},"_hasShrinkwrap":false}},"time":{"created":"2025-08-08T00:56:07.815Z","modified":"2025-08-09T05:38:22.734Z","0.3.9-alpha.0":"2025-08-08T00:56:08.167Z","0.3.9-alpha.1":"2025-08-09T05:38:22.515Z"},"author":{"name":"Daydreams"},"license":"MIT","keywords":["ai","agents","synthetic-data","training","machine-learning","fine-tuning"],"description":"Synthetic data generation for AI agent training","maintainers":[{"name":"realm_lord","email":"realm.lord.eth@gmail.com"},{"name":"bmgalego","email":"bmgalego@gmail.com"}],"readme":"# @daydreamsai/synthetic\n\nSynthetic data generation for AI agent training - creating a symbiotic\nrelationship between agent reasoning and model training.\n\n## Overview\n\nThis package captures your agent's reasoning process and converts it into\nhigh-quality training datasets. Turn on 'synthetic' generation and your agent\nautomatically generates perfectly formatted datasets from its thoughts, actions,\nand conversations.\n\n## Key Features\n\n- **Real-time data capture** - Monitor agent reasoning as it happens\n- **Multiple export formats** - Instruction tuning, conversation, reasoning\n  chains, action sequences, episodes, GRPO preference data\n- **Quality analysis** - Built-in quality scoring and issue detection\n- **Privacy controls** - Redact sensitive patterns, anonymize users\n- **Seamless integration** - Works with any Daydreams agent through extensions\n\n## Quick Start\n\n### Basic Usage\n\n```typescript\nimport { createDreams } from \"@daydreamsai/core\";\nimport { createSyntheticData } from \"@daydreamsai/synthetic\";\n\nconst agent = createDreams({\n  model: groq(\"deepseek-r1-distill-llama-70b\"),\n  contexts: [cli],\n  extensions: [\n    createSyntheticData({\n      enabled: true,\n      outputDir: \"./training-data\",\n      formats: [\"instruction-tuning\", \"conversation\"],\n      mode: \"realtime\",\n    }),\n  ],\n});\n```\n\n### Advanced Configuration\n\n```typescript\nimport { createSyntheticExtension } from \"@daydreamsai/synthetic\";\n\nconst syntheticExtension = createSyntheticExtension({\n  enabled: true,\n  outputDir: \"./synthetic-data\",\n  formats: [\"instruction-tuning\", \"reasoning-chains\", \"episodes\"],\n\n  capture: {\n    conversations: true,\n    reasoning: true,\n    actions: true,\n    episodes: true,\n  },\n\n  filters: {\n    minConversationLength: 3,\n    successfulOnly: false,\n    contexts: [\"cli\", \"discord\"],\n    actions: [\"web_search\", \"calculate\"],\n  },\n\n  privacy: {\n    redactPatterns: [/\\b[A-Za-z0-9._%+-]+@[A-Za-z0-9.-]+\\.[A-Z|a-z]{2,}\\b/],\n    anonymizeUsers: true,\n    removeTimestamps: false,\n  },\n\n  mode: \"batch\",\n  batchSize: 100,\n});\n```\n\n## Data Formats\n\n### Instruction Tuning Format\n\nPerfect for fine-tuning language models:\n\n```json\n{\n  \"instruction\": \"What's the weather like today?\",\n  \"response\": \"I can search for current weather information for you. What city would you like me to check?\",\n  \"system\": \"You are a helpful AI assistant agent.\",\n  \"context\": \"Context: cli, ID: session_123\"\n}\n```\n\n### Conversation Format\n\nMulti-turn dialogue training:\n\n```json\n{\n  \"messages\": [\n    { \"role\": \"user\", \"content\": \"Hello!\" },\n    { \"role\": \"assistant\", \"content\": \"Hi! How can I help you today?\" },\n    { \"role\": \"user\", \"content\": \"I need help with my tasks\" },\n    {\n      \"role\": \"assistant\",\n      \"content\": \"I'd be happy to help you organize your tasks. What would you like to work on?\"\n    }\n  ],\n  \"summary\": \"Conversation with 2 user messages and 2 assistant responses\"\n}\n```\n\n### Reasoning Chains Format\n\nStep-by-step thinking for chain-of-thought training:\n\n```json\n{\n  \"problem\": \"Calculate the ROI for a $10,000 investment that returned $12,500\",\n  \"reasoning\": [\n    {\n      \"step\": 1,\n      \"thought\": \"I need to calculate ROI using the formula (Returns - Investment) / Investment * 100\"\n    },\n    { \"step\": 2, \"thought\": \"Returns = $12,500, Investment = $10,000\" },\n    {\n      \"step\": 3,\n      \"action\": \"calculate({\\\"operation\\\": \\\"subtract\\\", \\\"a\\\": 12500, \\\"b\\\": 10000})\",\n      \"result\": \"2500\"\n    },\n    {\n      \"step\": 4,\n      \"action\": \"calculate({\\\"operation\\\": \\\"divide\\\", \\\"a\\\": 2500, \\\"b\\\": 10000})\",\n      \"result\": \"0.25\"\n    },\n    {\n      \"step\": 5,\n      \"action\": \"calculate({\\\"operation\\\": \\\"multiply\\\", \\\"a\\\": 0.25, \\\"b\\\": 100})\",\n      \"result\": \"25\"\n    }\n  ],\n  \"conclusion\": \"The ROI is 25%\"\n}\n```\n\n### Action Sequences Format\n\nAction usage patterns:\n\n```json\n{\n  \"situation\": \"User wants to know the weather\",\n  \"actions\": [\n    {\n      \"name\": \"web_search\",\n      \"arguments\": { \"query\": \"current weather San Francisco\" },\n      \"result\": { \"temperature\": \"72°F\", \"conditions\": \"sunny\" },\n      \"timestamp\": 1704067200000,\n      \"success\": true\n    }\n  ],\n  \"outcome\": \"Provided current weather information for San Francisco\"\n}\n```\n\n### Episodes Format\n\nComplete interaction episodes:\n\n```json\n{\n  \"episodeId\": \"session_123\",\n  \"observation\": \"User asked for help with productivity\",\n  \"thoughts\": [\n    \"The user needs help with productivity\",\n    \"I should suggest task management strategies\",\n    \"Let me provide specific actionable advice\"\n  ],\n  \"actions\": [\n    {\n      \"name\": \"suggest_tasks\",\n      \"arguments\": { \"category\": \"productivity\" },\n      \"result\": [\"Use time blocking\", \"Set priorities\", \"Take breaks\"],\n      \"timestamp\": 1704067200000,\n      \"success\": true\n    }\n  ],\n  \"result\": \"Provided productivity advice and task management suggestions\",\n  \"success\": true,\n  \"duration\": 15000\n}\n```\n\n### GRPO Format\n\nGroup Relative Policy Optimization training with preference data:\n\n```json\n{\n  \"prompt\": \"Explain the concept of machine learning in simple terms\",\n  \"responses\": [\n    {\n      \"text\": \"Machine learning is a type of artificial intelligence where computers learn to make predictions or decisions by finding patterns in data, rather than being explicitly programmed for every scenario.\",\n      \"score\": 0.9,\n      \"rank\": 1,\n      \"success\": true,\n      \"model\": \"gpt-4\",\n      \"metadata\": { \"outputId\": \"out_001\", \"timestamp\": 1704067705000 }\n    },\n    {\n      \"text\": \"Machine learning is when computers learn stuff from data to make predictions.\",\n      \"score\": 0.6,\n      \"rank\": 2,\n      \"success\": true,\n      \"model\": \"gpt-3.5\",\n      \"metadata\": { \"outputId\": \"out_002\", \"timestamp\": 1704067706000 }\n    },\n    {\n      \"text\": \"I'm not sure how to explain machine learning.\",\n      \"score\": 0.2,\n      \"rank\": 3,\n      \"success\": false,\n      \"model\": \"basic-model\",\n      \"metadata\": { \"outputId\": \"out_003\", \"timestamp\": 1704067707000 }\n    }\n  ],\n  \"system\": \"You are a helpful AI assistant agent.\",\n  \"context\": \"Context: cli, ID: session_011\",\n  \"comparisons\": [\n    { \"preferred\": 0, \"rejected\": 1, \"confidence\": 0.6 },\n    { \"preferred\": 0, \"rejected\": 2, \"confidence\": 0.9 },\n    { \"preferred\": 1, \"rejected\": 2, \"confidence\": 0.8 }\n  ]\n}\n```\n\n## Agent Actions\n\nThe extension adds these actions to your agent:\n\n### Process Data\n\n```typescript\n// Process accumulated logs and export\nawait agent.callAction(\"synthetic.process\", {\n  export: true,\n  analyze: true,\n});\n```\n\n### Configure Settings\n\n```typescript\n// Update configuration on the fly\nawait agent.callAction(\"synthetic.configure\", {\n  enabled: true,\n  mode: \"batch\",\n  formats: [\"instruction-tuning\", \"episodes\"],\n});\n```\n\n### Analyze Quality\n\n```typescript\n// Analyze data quality\nawait agent.callAction(\"synthetic.analyze\", {\n  filePath: \"./synthetic-data/dataset.jsonl\",\n});\n```\n\n### Export Episodes\n\n```typescript\n// Export all stored episodes\nawait agent.callAction(\"synthetic.exportAllEpisodes\", {\n  format: \"episodes\",\n});\n```\n\n## Quality Analysis\n\nThe package includes comprehensive quality analysis:\n\n```typescript\nimport { SyntheticAnalyzer } from \"@daydreamsai/synthetic\";\n\nconst analyzer = new SyntheticAnalyzer();\nconst quality = analyzer.analyzeQuality(records);\n\nconsole.log(quality);\n// {\n//   overallScore: 0.85,\n//   diversity: 0.92,\n//   completeness: 0.88,\n//   consistency: 0.75,\n//   byFormat: {\n//     \"instruction-tuning\": 0.90,\n//     \"conversation\": 0.85,\n//     \"reasoning-chains\": 0.80\n//   }\n// }\n```\n\n## Privacy & Security\n\nBuilt-in privacy controls:\n\n```typescript\nconst config = {\n  privacy: {\n    // Redact email addresses\n    redactPatterns: [/\\b[A-Za-z0-9._%+-]+@[A-Za-z0-9.-]+\\.[A-Z|a-z]{2,}\\b/],\n\n    // Anonymize user IDs\n    anonymizeUsers: true,\n\n    // Remove timestamps for privacy\n    removeTimestamps: false,\n  },\n};\n```\n\n## Use Cases\n\n### Fine-tuning Personal Assistant\n\nCapture your agent's conversations to create a personalized model:\n\n```typescript\ncreateSyntheticData({\n  formats: [\"instruction-tuning\", \"conversation\"],\n  capture: { conversations: true, reasoning: false },\n  filters: { minConversationLength: 2, successfulOnly: true },\n});\n```\n\n### Training Reasoning Models\n\nCapture step-by-step thinking for chain-of-thought models:\n\n```typescript\ncreateSyntheticData({\n  formats: [\"reasoning-chains\"],\n  capture: { reasoning: true, actions: true },\n  filters: { contexts: [\"complex-problem-solving\"] },\n});\n```\n\n### Action Pattern Learning\n\nLearn from successful action sequences:\n\n```typescript\ncreateSyntheticData({\n  formats: [\"action-sequences\", \"episodes\"],\n  capture: { actions: true, episodes: true },\n  filters: { successfulOnly: true },\n});\n```\n\n### GRPO Preference Training\n\nGenerate preference data for Group Relative Policy Optimization:\n\n```typescript\ncreateSyntheticData({\n  formats: [\"grpo\"],\n  capture: { preferences: true },\n  filters: {\n    minConversationLength: 1,\n    successfulOnly: false, // Include failures for better preference learning\n  },\n});\n```\n\n## Integration with Training\n\nExport data in formats ready for popular training frameworks:\n\n### Hugging Face Transformers\n\n```bash\n# JSONL format ready for transformers\npython train.py --data_path ./synthetic-data/instruction-tuning-*.jsonl\n```\n\n### OpenAI Fine-tuning\n\n```bash\n# Upload instruction tuning data\nopenai api fine_tunes.create -t ./synthetic-data/instruction-tuning-*.jsonl -m davinci\n```\n\n### Custom Training Scripts\n\n```python\nimport json\n\n# Load synthetic conversation data\nwith open('./synthetic-data/conversation-*.jsonl', 'r') as f:\n    for line in f:\n        record = json.loads(line)\n        messages = record['messages']\n        # Process for your training pipeline\n```\n\n## Best Practices\n\n### 1. Start Small\n\nBegin with basic conversation capture:\n\n```typescript\ncreateSyntheticData({\n  formats: [\"instruction-tuning\"],\n  mode: \"batch\",\n  batchSize: 10,\n});\n```\n\n### 2. Filter Quality Data\n\nUse filters to ensure high-quality training data:\n\n```typescript\n{\n  filters: {\n    minConversationLength: 3,\n    successfulOnly: true,\n    contexts: [\"production-contexts\"]\n  }\n}\n```\n\n### 3. Monitor Quality\n\nRegularly analyze your synthetic data:\n\n```typescript\n// Check quality metrics\nawait agent.callAction(\"synthetic.analyze\");\n\n// Export and review\nawait agent.callAction(\"synthetic.process\", { export: true });\n```\n\n### 4. Privacy First\n\nAlways configure privacy controls:\n\n```typescript\n{\n  privacy: {\n    redactPatterns: [/sensitive-pattern/g],\n    anonymizeUsers: true\n  }\n}\n```\n\n## API Reference\n\n### SyntheticConfig\n\nComplete configuration interface for synthetic data generation.\n\n### SyntheticRecord\n\nIndividual training record with metadata and quality scoring.\n\n### RealtimeSyntheticCollector\n\nProcesses agent logs in real-time into training records.\n\n### SyntheticExporter\n\nExports records to various file formats (JSONL, JSON).\n\n### SyntheticAnalyzer\n\nAnalyzes data quality and detects issues.\n\n## Development\n\n### Building\n\n```bash\nnpm run build\n```\n\n### Testing\n\n```bash\nnpm test\n```\n\n### Development Mode\n\n```bash\nnpm run dev\n```\n\n## License\n\nMIT\n\n## Contributing\n\nContributions welcome! Please read our contributing guidelines and submit pull\nrequests to our repository.\n","readmeFilename":"README.md"}