{"_id":"@biks2013/image-tool","name":"@biks2013/image-tool","dist-tags":{"latest":"0.1.0"},"versions":{"0.1.0":{"name":"@biks2013/image-tool","version":"0.1.0","description":"TypeScript CLI for generating, editing, examining, and varying images via Azure OpenAI and Google Gemini.","type":"module","license":"MIT","author":{"name":"Giorgos Marinos","email":"giorgos.marinos@gmail.com"},"repository":{"type":"git","url":"git+https://github.com/BikS2013/image-tool.git"},"homepage":"https://github.com/BikS2013/image-tool#readme","bugs":{"url":"https://github.com/BikS2013/image-tool/issues"},"keywords":["image","cli","azure-openai","gemini","image-generation","image-editing","langgraph","langchain","agent","tui","typescript"],"engines":{"node":">=20.0.0"},"bin":{"image-tool":"dist/bin/image-tool.js"},"main":"dist/index.js","publishConfig":{"access":"public"},"scripts":{"build":"tsc && chmod +x dist/bin/image-tool.js","dev":"tsx src/bin/image-tool.ts","start":"node dist/bin/image-tool.js","test":"vitest run","test:watch":"vitest","lint":"tsc --noEmit","format":"prettier --write .","typecheck":"tsc --noEmit","clean":"rm -rf dist","prepublishOnly":"npm run clean && npm run build && npm test"},"dependencies":{"@clack/prompts":"^0.11.0","@google/genai":"^1.0.0","@langchain/anthropic":"^1.3.27","@langchain/core":"^1.1.41","@langchain/google-genai":"^2.1.28","@langchain/langgraph":"^1.2.9","@langchain/openai":"^1.4.4","commander":"^14.0.0","consola":"^3.0.0","dotenv":"^17.0.0","openai":"^6.0.0","zod":"^4.0.0"},"devDependencies":{"@types/node":"^20.0.0","prettier":"^3.0.0","tsx":"^4.0.0","typescript":"^5.6.0","vitest":"^3.0.0"},"overrides":{"node-domexception":"file:./shims/node-domexception-shim"},"gitHead":"0b1651bfc01abdc82d3c4e397f0562e3249ada50","_id":"@biks2013/image-tool@0.1.0","_nodeVersion":"25.9.0","_npmVersion":"11.12.1","dist":{"integrity":"sha512-cGEgtdixiJlUi60xgvgqNWi6AXNSZwq/Dh075TvYwgjdPyRZHsb2hYCcrWsJ1wIr27oljFjiuE4+a7ga0CHSJQ==","shasum":"9cd83433da2c344b96eb9b9960c772b6ef16e979","tarball":"https://registry.npmjs.org/@biks2013/image-tool/-/image-tool-0.1.0.tgz","fileCount":279,"unpackedSize":793956,"signatures":[{"keyid":"SHA256:DhQ8wR5APBvFHLF/+Tc+AYvPOdTpcIDqOhxsBHRwC7U","sig":"MEUCIGTlFZFQthgmpx6ncuSgsfEMhT16smFHFomh/FWHBQ6HAiEA2NSPgX2jMNiX6eaRD5IV9LJsMFtbTziNAUKd1EWA8sc="}]},"_npmUser":{"name":"giorgos-marinos","email":"giorgos.marinos@gmail.com"},"directories":{},"maintainers":[{"name":"giorgos-marinos","email":"giorgos.marinos@gmail.com"}],"_npmOperationalInternal":{"host":"s3://npm-registry-packages-npm-production","tmp":"tmp/image-tool_0.1.0_1777188447537_0.5142758366309126"},"_hasShrinkwrap":false}},"time":{"created":"2026-04-26T07:27:27.437Z","0.1.0":"2026-04-26T07:27:27.677Z","modified":"2026-04-26T07:27:27.928Z"},"maintainers":[{"name":"giorgos-marinos","email":"giorgos.marinos@gmail.com"}],"description":"TypeScript CLI for generating, editing, examining, and varying images via Azure OpenAI and Google Gemini.","homepage":"https://github.com/BikS2013/image-tool#readme","keywords":["image","cli","azure-openai","gemini","image-generation","image-editing","langgraph","langchain","agent","tui","typescript"],"repository":{"type":"git","url":"git+https://github.com/BikS2013/image-tool.git"},"author":{"name":"Giorgos Marinos","email":"giorgos.marinos@gmail.com"},"bugs":{"url":"https://github.com/BikS2013/image-tool/issues"},"license":"MIT","readme":"# image-tool\n\nA TypeScript CLI for generating, editing, examining, and varying images via Azure OpenAI and Google Gemini.\n\n## Install\n\n```sh\nnpm install\n```\n\n## Build\n\n```sh\nnpm run build\n```\n\n## Dev\n\n```sh\nnpm run dev -- generate --provider gemini --prompt \"a cat\" --out cat.png\n```\n\n## Quick start\n\n```sh\ncp .env.example .env\n# edit .env and fill in credentials for the provider(s) you want to use\nnpm install\nnpm run build\n./dist/bin/image-tool.js generate --provider gemini --prompt \"a cat\" --out cat.png\n```\n\nFor the full functional specification (subcommands, flags, error model, exit codes, configuration contract, capability matrix), see [`docs/design/refined-request-image-tool-cli.md`](docs/design/refined-request-image-tool-cli.md).\n\n## Configuration\n\nAll configuration is supplied through environment variables (loaded from `.env` via dotenv). No defaults are applied except for `IMAGE_TOOL_MAX_CONCURRENCY`. Missing required values raise `MissingConfigurationError` (exit code 2).\n\n| Variable | Description |\n|---|---|\n| `IMAGE_TOOL_DEFAULT_PROVIDER` | Default provider (`openai-azure` \\| `gemini`) when neither `--provider` nor a REPL session override is supplied. |\n| `IMAGE_TOOL_DEFAULT_MODEL` | Default model identifier resolved when `--model` is not supplied. |\n| `IMAGE_TOOL_MAX_CONCURRENCY` | Cap on parallel provider calls inside one command. Optional positive integer. Default: `4` (the only sanctioned default). |\n| `IMAGE_TOOL_OUTPUT_DIR` | Optional output directory for generated artifacts. Defaults to the current working directory. |\n| `IMAGE_TOOL_LOG_LEVEL` | Optional log verbosity override (`silent` \\| `error` \\| `warn` \\| `info` \\| `debug` \\| `trace`). |\n| `AZURE_OPENAI_API_KEY` | Azure OpenAI API key. Required when provider = `openai-azure`. |\n| `AZURE_OPENAI_ENDPOINT` | Azure OpenAI resource endpoint URL. Required when provider = `openai-azure`. |\n| `AZURE_OPENAI_DEPLOYMENT_NAME` | Azure deployment name for the image-generation model. Required when provider = `openai-azure` for `generate` / `edit` / `variations`. |\n| `AZURE_OPENAI_API_VERSION` | Azure OpenAI REST API version supporting the deployed image model. Required when provider = `openai-azure`. |\n| `AZURE_OPENAI_VISION_DEPLOYMENT_NAME` | Azure deployment name for a vision-capable chat model. Required when provider = `openai-azure` and operation = `examine`. |\n| `GOOGLE_API_KEY` | Google AI Studio API key (preferred). Required when provider = `gemini` if `GEMINI_API_KEY` is not set. |\n| `GEMINI_API_KEY` | Google AI Studio API key (fallback alias). Required when provider = `gemini` if `GOOGLE_API_KEY` is not set. |\n| `GEMINI_IMAGE_MODEL` | Gemini model identifier for image operations (e.g. `gemini-3.1-flash-image-preview`). Required when provider = `gemini`. |\n\nSee `.env.example` for a complete annotated template.\n\n## Subcommands\n\nThe CLI exposes the following subcommands (implementation arrives in subsequent units):\n\n- `generate` — produce a new image from a text prompt.\n- `edit` — modify an existing image using a prompt and optional mask.\n- `examine` — describe / answer questions about an existing image.\n- `variations` — produce N variations of an existing image.\n- `config show` — print the resolved configuration (with secrets redacted).\n- `config doctor` — diagnose configuration problems and suggest fixes.\n- `repl` — open an interactive session that remembers provider / model selections.\n\nRun `image-tool <subcommand> --help` for per-command flags once the binary is built.\n\n## Agent mode\n\n`image-tool agent` adds a LangGraph ReAct agent that drives the four image operations\nthrough an LLM. The LLM decides which tools to call, chains them (generate → examine → edit),\nand reports results in plain prose.\n\n### Quick start\n\n```bash\n# Set the LLM provider and provider-specific credentials\nexport IMAGE_TOOL_AGENT_PROVIDER=azure-openai\nexport AZURE_OPENAI_AGENT_DEPLOYMENT=gpt-4.1   # your chat deployment\n\n# One-shot\nimage-tool agent \"generate a watercolor sunset and describe what you see\"\n\n# Interactive REPL\nimage-tool agent --interactive\n\n# Inspect resolved configuration (no LLM calls)\nimage-tool agent --doctor\n```\n\n### Worked examples\n\n**1. Azure OpenAI one-shot**\n```bash\nexport IMAGE_TOOL_AGENT_PROVIDER=azure-openai\nexport AZURE_OPENAI_AGENT_DEPLOYMENT=gpt-4.1\nimage-tool agent \"generate a futuristic city at dusk, then tell me the dominant colours\"\n```\n\n**2. Azure OpenAI interactive REPL**\n```bash\nimage-tool agent --llm-provider azure-openai --interactive\n# > generate a kitten playing with yarn\n# > now make it look more like a watercolour painting\n# > /exit\n```\n\n**3. JSON mode (CI-friendly)**\n```bash\nimage-tool agent --json --no-stream \"generate a minimalist logo for a coffee shop\"\n```\n\n**4. Local model with OLLaMA**\n```bash\nexport OLLAMA_HOST=http://localhost:11434\nexport LOCAL_LLM_MODEL=llama3.1\nimage-tool agent --llm-provider local-openai-compatible \"what image formats can you generate?\"\n```\n\n### Provider setup\n\n| Provider id | Required env vars |\n|---|---|\n| `azure-openai` | `AZURE_OPENAI_API_KEY`, `AZURE_OPENAI_ENDPOINT`, `AZURE_OPENAI_AGENT_DEPLOYMENT` |\n| `openai` | `OPENAI_API_KEY`, `OPENAI_AGENT_MODEL` |\n| `anthropic` | `ANTHROPIC_API_KEY`, `ANTHROPIC_AGENT_MODEL` |\n| `gemini` | `GOOGLE_API_KEY` (or `GEMINI_API_KEY`), `GEMINI_AGENT_MODEL` |\n| `azure-anthropic` | `AZURE_ANTHROPIC_API_KEY`, `AZURE_ANTHROPIC_ENDPOINT`, `AZURE_ANTHROPIC_DEPLOYMENT`, `AZURE_ANTHROPIC_API_VERSION` |\n| `local-openai-compatible` | `LOCAL_LLM_BASE_URL` (or `OLLAMA_HOST`), `LOCAL_LLM_MODEL` |\n\nSee `docs/reference/.env.example` for a fully annotated template.\n\n### Configuration precedence\n\nPolicy A (shell-wins):\n```\nCLI flag  >  shell env var  >  ~/.tool-agents/image-tool/config  >  project .env  >  error\n```\n\n### Safety note\n\nAll four image tools are **auto-approved** — the agent invokes them without per-call\nconfirmation. Images land in `.image-tool-agent/<timestamp>/` under the current working\ndirectory unless you specify `outputDir` in your prompt.\n\n### TUI (interactive mode)\n\nRunning `image-tool agent` with no positional prompt opens a raw-mode terminal UI on top of\nthe same LangGraph agent. The TUI gives you multi-line editing, token-by-token streaming,\nESC-to-abort, persistent JSONL session transcripts, and a small set of slash commands.\n\n```bash\n# Launch\nimage-tool agent\n```\n\n**Slash commands** — `/help`, `/history [N]`, `/memory`, `/new`, `/last`, `/copy`,\n`/model <id>`, `/provider <id>`, `/quit` (alias `/exit`).\n\n**Keybindings**\n\n| Key | Action |\n|---|---|\n| Enter | Submit |\n| Ctrl-J or Alt-Enter | Insert newline (universal Shift+Enter fallback) |\n| Up / Down | Browse user-input history (or move row in multi-line input) |\n| Left / Right | Cursor by character |\n| Home / End / Ctrl-A / Ctrl-E | Line start / end |\n| Alt-←/→ or Ctrl-←/→ | Word motion |\n| Ctrl-W / Ctrl-U / Ctrl-K | Delete word back / to start / to end |\n| Backspace / Delete | Delete left / delete at cursor |\n| ESC or Ctrl-C while streaming | Abort the in-flight agent response |\n| Ctrl-D on empty buffer | Quit cleanly |\n\n**About Shift+Enter** — most terminals (Apple Terminal, default iTerm2, WezTerm) send\nplain `\\r` for both Enter and Shift+Enter, so they are indistinguishable. The TUI accepts\nthe modern keyboard-protocol variants emitted by Kitty, Ghostty, Alacritty, Windows\nTerminal, and xterm with `modifyOtherKeys=2`. The universal portable fallback is\n**Ctrl-J** (literal LF byte 0x0A) which every terminal emits unambiguously.\n\n**Persistence** — each session writes a JSONL transcript under\n`~/.tool-agents/image-tool/history/${ISO8601}-${shortid}.jsonl` (directory mode 0700,\nfiles mode 0600). Override the directory with the optional env var\n`IMAGE_TOOL_TUI_HISTORY_DIR`.\n\n**Example transcript**\n\n```\nimage-tool agent (LangGraph)\nLLM: openai / gpt-4.1\nSession: kx7m2a → /Users/me/.tool-agents/image-tool/history/2026-04-24T20-15-03-kx7m2a.jsonl\nCommands: /help /history /memory /new /last /copy /model /provider /quit\nKeys: Enter=submit · Ctrl-J or Alt-Enter=newline · Up/Down=history · ESC=abort streaming · Ctrl-D=quit\n\n❯ generate a watercolor mountain at dusk\n\nAgent\n  ▸ generate_image  {\"prompt\":\"a watercolor mountain at dusk\"} › ✓\n  · /Users/me/work/.image-tool-agent/20260424-201510/img-001.png\nI have generated the watercolor mountain at dusk for you.\n[openai · gpt-4.1 · 1 turns · /help]\n\n❯ /quit\n```\n\n## Filesystem operations\n\n`image-tool fs <subcommand>` exposes 10 first-class filesystem helpers\nthat are also available as LangChain tools to the agent. The CLI side is\nNOT sandboxed (paths resolve relative to the current directory and the\nhuman user is trusted); the agent side enforces a sandbox rooted at\n`IMAGE_TOOL_FS_ROOT` (or `process.cwd()` when unset) and gates the\ndestructive ops behind an interactive confirmation prompt.\n\n```\nimage-tool fs ls /tmp [-a] [--json]\nimage-tool fs stat /etc/hosts [--json]\nimage-tool fs read /etc/hosts [--max-bytes 4096] [--offset 0] [--base64-only] [--json]\nimage-tool fs write /tmp/x.txt --content \"hello\" [--overwrite] [--no-mkdir-p]\nimage-tool fs append /tmp/x.txt --from-stdin\nimage-tool fs mkdir -p /tmp/a/b/c\nimage-tool fs rm -rf /tmp/scratch\nimage-tool fs mv /tmp/old.txt /tmp/new.txt --overwrite\nimage-tool fs cp -r /tmp/src /tmp/dst\nimage-tool fs find \"**/*.json\" --root . --max-results 100\n```\n\nWhen `--json` is passed the command emits a single JSON object on stdout\nthat mirrors the typed result in `src/fs/types.ts` — useful for scripting.\n\nThe agent catalog grows from 4 image tools to 14 once the fs tools are\nincluded (10 fs + 4 image). Destructive agent ops (`fs_rm`, `fs_mv`,\n`fs_cp`) require interactive confirmation; in one-shot mode they\nreturn a typed `FsConfirmationRequiredError` to the LLM so the agent\nsees a clear \"the user needs to switch to interactive mode\" signal\ninstead of silently failing.\n\n### Optional environment\n\n```\nIMAGE_TOOL_FS_ROOT     Absolute sandbox root for the agent fs tools.\n                       Defaults to `process.cwd()` at agent launch.\n                       Has no effect on the CLI subcommands.\n```\n\n","readmeFilename":"README.md","_rev":"1-23ae3ed53697e1f1e328b8c73f94f854"}