{"_id":"@aksr333/jarvis","_rev":"2-538ac74ef5165e633faad99d5279f0f9","name":"@aksr333/jarvis","dist-tags":{"latest":"1.0.1"},"versions":{"1.0.0":{"name":"@aksr333/jarvis","version":"1.0.0","keywords":["jarvis","voice-assistant","ai","macos","speech-to-text","llm","cerebras","deepgram","ollama","whisper"],"author":{"name":"Akshath"},"license":"MIT","_id":"@aksr333/jarvis@1.0.0","maintainers":[{"name":"aksr333","email":"akshath.r333@gmail.com"}],"homepage":"https://github.com/akshath-raj/JARVIS#readme","bugs":{"url":"https://github.com/akshath-raj/JARVIS/issues"},"os":["darwin"],"bin":{"jarvis":"bin/jarvis.js"},"dist":{"shasum":"da4990f3ac9da0b8ca058c6b7a157bd1c9ea4e3c","tarball":"https://registry.npmjs.org/@aksr333/jarvis/-/jarvis-1.0.0.tgz","fileCount":96,"integrity":"sha512-c83lPpcZ/WhvLvd5XjVTZg/tMDCGYYndz960Kt8W8FVqPN2zx0LQkoc5DGvG8cUREsRKxbKiPf2nT00EzZhXlQ==","signatures":[{"sig":"MEUCIQD4KJpfFAeQsVg+vCZHrZmOR+4em+2sd87EG5VFaEZEDwIgKBRfQoJAE2QvL/kPiZWxecATrUGE1kIgPcMTb1qh9aQ=","keyid":"SHA256:DhQ8wR5APBvFHLF/+Tc+AYvPOdTpcIDqOhxsBHRwC7U"}],"unpackedSize":597728},"type":"commonjs","engines":{"node":">=18"},"gitHead":"bf5a776635bc8ff6cba37e817d50946e26a18973","scripts":{"prepack":"find jarvis scripts -name '__pycache__' -type d -prune -exec rm -rf {} + 2>/dev/null || true","postinstall":"node bin/jarvis.js postinstall"},"_npmUser":{"name":"aksr333","email":"akshath.r333@gmail.com"},"repository":{"url":"git+https://github.com/akshath-raj/JARVIS.git","type":"git"},"_npmVersion":"11.9.0","description":"Iron-Man-style, voice-controlled AI assistant for macOS — plays music, drives the browser, reads your screen, answers from your documents, and more. One `jarvis` command sets everything up.","directories":{},"_nodeVersion":"25.6.1","preferGlobal":true,"publishConfig":{"access":"public"},"_hasShrinkwrap":false,"_npmOperationalInternal":{"tmp":"tmp/jarvis_1.0.0_1786776731355_0.6301700909712105","host":"s3://npm-registry-packages-npm-production"}},"1.0.1":{"name":"@aksr333/jarvis","version":"1.0.1","description":"Iron-Man-style, voice-controlled AI assistant for macOS — plays music, drives the browser, reads your screen, answers from your documents, and more. One `jarvis` command sets everything up.","keywords":["jarvis","voice-assistant","ai","macos","speech-to-text","llm","cerebras","deepgram","ollama","whisper"],"homepage":"https://github.com/akshath-raj/JARVIS#readme","bugs":{"url":"https://github.com/akshath-raj/JARVIS/issues"},"repository":{"type":"git","url":"git+https://github.com/akshath-raj/JARVIS.git"},"license":"MIT","author":{"name":"Akshath"},"publishConfig":{"access":"public"},"type":"commonjs","bin":{"jarvis":"bin/jarvis.js"},"engines":{"node":">=18"},"os":["darwin"],"preferGlobal":true,"scripts":{"prepack":"find jarvis scripts -name '__pycache__' -type d -prune -exec rm -rf {} + 2>/dev/null || true","postinstall":"node bin/jarvis.js postinstall"},"gitHead":"0ed43795a30661a3ae6b13ccf04b5378d634c315","_id":"@aksr333/jarvis@1.0.1","_nodeVersion":"25.6.1","_npmVersion":"11.9.0","dist":{"integrity":"sha512-bkYy3XhIabr72/VjWK97Hkx0rwAbpWwRGWjepaeLbE4hjy167lRP3z0ltmNfkA00WXs/U68ZKv4GsoEWfTYTcw==","shasum":"8c573a3aa3a9e48b9f0bfd9250df9ffcabb6b70e","tarball":"https://registry.npmjs.org/@aksr333/jarvis/-/jarvis-1.0.1.tgz","fileCount":96,"unpackedSize":598468,"signatures":[{"keyid":"SHA256:DhQ8wR5APBvFHLF/+Tc+AYvPOdTpcIDqOhxsBHRwC7U","sig":"MEUCIQCXBlPPr+aTOgr+RHVbNyT3vWtya03tp/rBX95gF7N86gIgcqSwhxflmtXDVBKxfVmhz0OoA4e8kynTJXVimTZZShU="}]},"_npmUser":{"name":"aksr333","email":"akshath.r333@gmail.com"},"directories":{},"maintainers":[{"name":"aksr333","email":"akshath.r333@gmail.com"}],"_npmOperationalInternal":{"host":"s3://npm-registry-packages-npm-production","tmp":"tmp/jarvis_1.0.1_1786777163725_0.9577845478789675"},"_hasShrinkwrap":false}},"time":{"created":"2026-08-15T06:52:11.164Z","modified":"2026-08-15T06:59:24.035Z","1.0.0":"2026-08-15T06:52:11.503Z","1.0.1":"2026-08-15T06:59:23.873Z"},"bugs":{"url":"https://github.com/akshath-raj/JARVIS/issues"},"author":{"name":"Akshath"},"license":"MIT","homepage":"https://github.com/akshath-raj/JARVIS#readme","keywords":["jarvis","voice-assistant","ai","macos","speech-to-text","llm","cerebras","deepgram","ollama","whisper"],"repository":{"type":"git","url":"git+https://github.com/akshath-raj/JARVIS.git"},"description":"Iron-Man-style, voice-controlled AI assistant for macOS — plays music, drives the browser, reads your screen, answers from your documents, and more. One `jarvis` command sets everything up.","maintainers":[{"name":"aksr333","email":"akshath.r333@gmail.com"}],"readme":"# JARVIS — a voice AI assistant for macOS\n\n![Platform: macOS](https://img.shields.io/badge/platform-macOS-black?logo=apple)\n![Python 3.11+](https://img.shields.io/badge/python-3.11%2B-blue?logo=python&logoColor=white)\n![License: MIT](https://img.shields.io/badge/license-MIT-green)\n![Local-first](https://img.shields.io/badge/local--first-Ollama%20%2B%20Whisper%20%2B%20Kokoro-orange)\n\nA voice-enabled assistant in the spirit of Iron Man's JARVIS. Wake it with\n**\"Hey Jarvis,\"** at the beginning of a turn, then talk to it, and it controls your Mac: plays music, drives the\nbrowser, reads and explains what's on your screen, answers questions from your own\ndocuments, manages a calendar and reminders, organises your files, and runs a\nfocus-assist mode that closes distractions — all narrated back to you, with a\nhidden **Iron-Man-style HUD** that appears on launch.\n\nIt runs **cloud-first by default** (fast, hosted models) and can also run **fully\non-device**. One agent holds every tool and chains them for multi-step requests\n(*\"explain this graph and check the 2026 numbers\"* → looks at your screen, then\nsearches the web, then answers).\n\n> **Platform:** macOS only (Apple Silicon recommended). See\n> [Platform support](#platform-support) for why, and Windows/Linux notes.\n\n---\n\n## Contents\n- [Feature overview](#feature-overview)\n- [Two pipelines (`JARVIS_MODE`)](#two-pipelines-jarvis_mode)\n- [Quick install (npm)](#quick-install-npm)\n- [Install & run — step by step (manual)](#install--run--step-by-step-manual)\n- [Optional features & their setup](#optional-features--their-setup)\n- [Platform support](#platform-support)\n- [Configuration reference](#configuration-reference)\n- [Architecture](#architecture)\n- [Privacy & safety](#privacy--safety)\n- [Troubleshooting](#troubleshooting)\n\n---\n\n## Feature overview\n\nEverything below works with the default `openai` brain (a single agent on\nCerebras/OpenAI). Say the wake word, then your request.\n\n### 🎵 Music (Spotify)\nPlayback runs **in the background** (launched hidden, driven by AppleScript) so it\nnever steals focus. Needs the Spotify desktop app installed + logged in.\n- *\"play Bohemian Rhapsody by Queen\"*, *\"pause\"*, *\"skip\"*, *\"louder\"*, *\"what's playing?\"*\n- *\"loop this\"*, *\"play Weightless on repeat\"*\n- Playlists — *\"what playlists do I have\"*, *\"play my Focus playlist\"*, *\"add this to Favourites\"*\n- Your library — *\"play my most-played song\"*, *\"my top songs/artists\"*, *\"my liked songs\"*\n\n### 🌐 Browser & web (Chrome)\n- *\"open instagram\"*, *\"play some lofi on youtube\"*, *\"open a new MrBeast video\"*, *\"show me reels\"*\n- Live info via **Tavily** web search — *\"latest news on X\"*, *\"price of bitcoin\"*, *\"weather tomorrow\"*\n- **Browser agent** (logged-in, multi-step) — *\"download my OS notes from VTOP\"*, *\"check my AWS balance\"*\n- **Assignment workflow** — *\"download the latest assignment and explain it\"* → *\"finish it as a notebook\"* → *\"open it\"*\n\n### 🖥️ Screen vision & on-screen document help\n- *\"what's on my screen?\"*, *\"explain this error\"*, *\"read this for me\"* → screenshots and explains.\n- *\"explain this topic, I don't understand it\"* → knows **which document is live on\n  screen**, the **page/slide** you're on, and reads the **surrounding pages of that\n  subtopic** from your indexed files to explain it properly.\n\n### 📄 Your documents (local RAG)\nJARVIS indexes your PDFs / Word / PowerPoint / notes into a **local, self-updating**\nvector store. Ask by **description — no exact filename needed**:\n- *\"what does my ML assignment say about backprop?\"*, *\"explain page 4 of my calculus notes\"*\n- *\"summarise my thermodynamics notes\"*, *\"open my OS notes\"*, *\"where's my budget spreadsheet?\"*\n\n### 🗂️ Sandboxed file organiser\nRead / copy / move / open / tidy folders. **Never deletes** and is **confined to\nyour home folder**. *\"organise my downloads\"*, *\"move X to Y\"*, *\"open my recent downloads\"*.\n\n### 📅 Calendar, to-dos, reminders & alarms\n*\"remind me to submit at 9pm\"*, *\"remind me in an hour\"*, *\"add milk to my list\"*. A\ndue reminder **rings an alarm**, pops up on the HUD, and speaks; *\"stop the alarm\"* / *\"I'm done with X\"*.\n\n### 💡 Screen & sound controls\nJARVIS drives your Mac's own **display brightness**, **system volume**, and\n**appearance** — separate from a song's or a video's volume.\n- *\"set the brightness to 40%\"*, *\"dim the screen\"*, *\"brighter\"*\n- *\"turn the volume down\"*, *\"set the sound to 30%\"*, *\"mute\"* / *\"unmute\"*\n- *\"switch to dark mode\"* / *\"light mode\"*\n\n> Precise brightness needs the `brightness` CLI (`brew install brightness`);\n> without it JARVIS steps the brightness keys (needs **Accessibility** permission).\n\n### 📖 Reading mode\n*\"reading mode\"* / *\"set me up to read\"* turns your Mac into a comfortable reading\nspace — and *\"stop reading mode\"* / *\"I'm done reading\"* puts everything back:\n- **Warm tone** (Night Shift) so it's easy on the eyes,\n- a **dark, low-glare background** (Dark Mode),\n- **softer brightness** tuned for reading, and\n- **soft instrumental music** on Spotify in the background at a low volume.\n\nAll the settings it changes are **saved first and restored on exit**. Tune the\ndefaults with `JARVIS_READING_*` (see [Configuration](#configuration-reference)).\nFor the warm tone, install the `nightlight` CLI (`brew install nightlight`).\n\n### 🎯 Focus assist mode\n*\"focus mode\"*, *\"start a pomodoro\"* → **closes every open distraction** (Instagram,\nYouTube, Netflix, TikTok, Reddit, X…), starts a timer, and **keeps closing anything\nyou reopen** until *\"end focus mode\"*. Techniques: **pomodoro / classic / 52-17 /\n90-20 / flowtime**. Shown live on the HUD with a countdown.\n\n### 🟦 The HUD dashboard\nA hidden Iron-Man-style interface that **opens automatically on launch** (⌄ HIDE to\ndismiss; say *\"show the dashboard\"* to bring it back). Shows your profile/memories,\npast conversations, agenda, a **live Q&A feed**, **rendered explanations**, the\n**focus timer**, and **alarm pop-ups**. Local only.\n\n### 🧠 Persistent memory (local)\nLearns what you like and recalls it to personalise replies. *\"remember that I…\"*,\n*\"what do you know about me?\"*, *\"forget that…\"*. Stored in `~/.jarvis/memory/`.\n\n---\n\n## Two pipelines (`JARVIS_MODE`)\n\n| | `1` CLOUD (default) | `0` LOCAL |\n|---|---|---|\n| **STT** | Deepgram `nova-3` (streaming) | Whisper mlx (`small.en`) |\n| **Brain** | Cerebras `gpt-oss-120b` → OpenAI | Ollama `qwen2.5:7b-instruct` |\n| **Vision** | Cerebras `gemma-4-31b` → gpt-4o | (uses the cloud vision key) |\n| **TTS** | Deepgram Aura-2 (`aura-2-draco-en`) | Kokoro-82M |\n\n**CLOUD** is fast and hosted. **LOCAL** keeps the voice loop fully on-device.\nDocument RAG and long-term memory always use **local** Ollama embeddings in either mode.\n\n---\n\n## Quick install (npm)\n\nIf you just want to run it, install the launcher from npm. It bundles JARVIS,\nbuilds an isolated Python environment on first run, and creates your config — you\nonly need Python, Node, and (for the local models) Ollama on your Mac.\n\n```bash\n# one-time system tools (skip any you already have)\nbrew install python node ollama ffmpeg\nbrew services start ollama && ollama pull nomic-embed-text\n\n# install JARVIS\nnpm install -g @aksr333/jarvis\n\njarvis setup       # builds the Python env + creates ~/.jarvis/.env  (first run only)\njarvis config      # paste your Cerebras + Deepgram API keys (see Step 6 below)\njarvis             # start talking — say \"Hey Jarvis\"\n```\n\nThat's it. The `jarvis` command:\n\n| Command | What it does |\n|---|---|\n| `jarvis` | Start the voice assistant (auto-runs setup the first time) |\n| `jarvis setup` | Build the Python env, install deps, create your config |\n| `jarvis config` | Open `~/.jarvis/.env` to add/edit your API keys |\n| `jarvis doctor` | Check prerequisites and configuration |\n| `jarvis update` | Reinstall Python deps after upgrading the package |\n\nYour keys and Python environment live under `~/.jarvis/`, so they **survive\nreinstalls and upgrades**. Prefer not to install globally? Use\n`npx @aksr333/jarvis` instead of `jarvis`.\n\n> **Why a Python env?** JARVIS's voice + agent stack is Python; the npm package is a\n> thin launcher that sets that up for you. First `jarvis setup` downloads a few\n> hundred MB of ML deps and takes a few minutes.\n\n---\n\n## Install & run — step by step (manual)\n\n> Prefer to run from a clone, or want to see exactly what the launcher does?\n> **Follow these in order on a Mac.** ~15 minutes. **Steps 1–8 give you a talking\n> assistant** with document Q&A, focus mode, and the HUD. Everything else is in\n> [Optional features](#optional-features--their-setup). Every command is copy-paste.\n\n### Step 1 — Install Homebrew (skip if you already have `brew`)\nHomebrew is the macOS package manager. Paste this into **Terminal** (⌘-Space → \"Terminal\"):\n```bash\n/bin/bash -c \"$(curl -fsSL https://raw.githubusercontent.com/Homebrew/install/HEAD/install.sh)\"\n```\nWhen it finishes, run the two `echo`/`eval` lines it prints (to add `brew` to your PATH), or just close and reopen Terminal.\n\n### Step 2 — Install the system tools\n```bash\nbrew install python git ollama ffmpeg node\n```\n\n### Step 3 — Get the code\n```bash\ngit clone https://github.com/akshath-raj/JARVIS.git\ncd JARVIS\n```\n\n### Step 4 — Create the Python environment\n```bash\npython3 -m venv .venv\nsource .venv/bin/activate\npip install -r requirements.txt\n```\n> You'll run JARVIS from this folder with `.venv` active. To reactivate later:\n> `cd JARVIS && source .venv/bin/activate`.\n\n### Step 5 — Start Ollama and pull the embedding model\nDocument search and memory always run locally through Ollama, so this is required\neven in cloud mode:\n```bash\nbrew services start ollama      # runs Ollama in the background, and on login\nollama pull nomic-embed-text\n```\n\n### Step 6 — Get your two free API keys\nYou need two keys for the default (cloud) setup:\n\n| Key | What it powers | Where to get it |\n|---|---|---|\n| **Cerebras** | the brain, document answers, screen vision | 1. Sign up at **[cloud.cerebras.ai](https://cloud.cerebras.ai)** → 2. **API Keys** → **Create**. It starts with `csk-`. |\n| **Deepgram** | speech-to-text + text-to-speech (the voice) | 1. Sign up at **[console.deepgram.com/signup](https://console.deepgram.com/signup)** (free credit) → 2. **Create API Key**. |\n\n### Step 7 — Create your `.env` and paste the keys\n```bash\ncp .env.example .env\nopen -e .env            # opens the file in TextEdit\n```\nIn the file that opens, find these two lines and paste your keys after the `=`:\n```ini\nCEREBRAS_API_KEY=csk-paste-your-cerebras-key-here\nDEEPGRAM_API_KEY=paste-your-deepgram-key-here\n```\n**Save (⌘S) and close.** That's the minimum — everything else in `.env` already has\nsensible defaults. (Your `.env` holds secrets and is git-ignored; never share it.)\n\n### Step 8 — Run it\n```bash\npython -m jarvis.agent console\n```\n- The **HUD dashboard opens automatically** and boots with a \"Welcome back\" animation.\n- macOS pops up a **Microphone** permission request — click **Allow**.\n- Say: **\"Hey Jarvis, what's the tallest mountain in the world?\"**\n\nPrefer no terminal? Double-click **`scripts/Start JARVIS.command`** in Finder (first\ntime: right-click ▸ **Open** ▸ **Open** to get past Gatekeeper). Or build a real app\nwith `bash scripts/make_app.sh` → `~/Applications/JARVIS.app` (drag it to your Dock).\n\n### Step 9 — Grant permissions as features ask for them\nOpen **System Settings ▸ Privacy & Security** and add your **Terminal** (or\n`JARVIS.app`) under each of these. Most are also prompted automatically the first\ntime you use the feature:\n\n| Permission | Needed for |\n|---|---|\n| **Microphone** | hearing you (prompted on first run) |\n| **Screen Recording** | *\"explain my screen\"* / on-screen document help |\n| **Automation** | controlling Chrome & System Events — music, browser, **focus mode** |\n| **Accessibility** | reading the focused window's title (which document you mean) |\n\n> After granting a permission, **quit and reopen** the terminal/app so it takes effect.\n\n**You're done.** Try: *\"Play some lofi on YouTube.\"* · *\"Focus mode.\"* ·\n*\"What's on my screen?\"* · *\"Remind me to stretch in 20 minutes.\"* ·\n*\"Summarise my notes.\"* (put a PDF in `~/Downloads` first).\n\n---\n\n## Optional features & their setup\n\nAdd any of these later — each is independent. Put keys in the same `.env`.\n\n<details>\n<summary><b>🎵 Spotify (music)</b></summary>\n\n1. Install the **Spotify desktop app** and log in (Premium recommended).\n2. Create a free app at **[developer.spotify.com/dashboard](https://developer.spotify.com/dashboard)**, copy the **Client ID** and **Client Secret** into `.env`:\n   ```ini\n   SPOTIFY_CLIENT_ID=...\n   SPOTIFY_CLIENT_SECRET=...\n   ```\n3. **Playlists & library** (not just search) need a one-time login. In the Spotify\n   dashboard, add `http://127.0.0.1:8080/callback` to your app's **Redirect URIs**, then:\n   ```bash\n   python -m jarvis.spotify_auth\n   ```\n</details>\n\n<details>\n<summary><b>🌐 Tavily (live web search)</b></summary>\n\nGet a free key at **[app.tavily.com](https://app.tavily.com)** and add:\n```ini\nTAVILY_API_KEY=tvly-...\n```\nEnables *\"what's the latest news on…\"*, *\"current price of…\"*, *\"weather tomorrow\"*.\n</details>\n\n<details>\n<summary><b>🤖 OpenAI (fallback brain/vision, browser agent, assignments)</b></summary>\n\nGet a key at **[platform.openai.com/api-keys](https://platform.openai.com/api-keys)**:\n```ini\nOPENAI_API_KEY=sk-...\n```\nUsed as a fallback if Cerebras is unavailable, and required by the browser agent\nand the assignment workflow (below).\n</details>\n\n<details>\n<summary><b>🧭 Browser agent (logged-in web workflows)</b></summary>\n\nMulti-step tasks on sites you're logged into (*\"download my notes from VTOP\"*). By\ndefault JARVIS uses **Playwright MCP** through Node.js (installed in Step 2). It opens\na visible, dedicated Chrome profile at `~/.jarvis/pw-profile`; sign in there once and\nthe session is retained for future tasks. It reads accessibility snapshots, uses vision\nwhen needed for custom/canvas controls, and handles normal multi-tab navigation,\nuploads, downloads, and dynamic pages.\n\nNeeds `OPENAI_API_KEY`. It acts only on spoken browser-task requests and stops for a\nlogin, MFA, CAPTCHA, paywall, or other human check rather than trying to bypass it.\n\nThe older `browser-use` runner is available only as a fallback when\n`JARVIS_PW_AGENT=0`. It runs in its own venv because its dependencies conflict with\nthe main stack:\n\n```bash\nbash scripts/setup_browser_agent.sh\n# optional local captcha solver:\nollama pull qwen2.5vl:7b\n```\n\n⚠️ Browsed page content is supplied to the browser model, so do not use it for a task\nthat exposes secrets you do not want the model to see. Page text is treated as data,\nnot instructions.\n</details>\n\n<details>\n<summary><b>🔒 Fully local mode (no cloud, nothing leaves your Mac)</b></summary>\n\nSet `JARVIS_MODE=0` in `.env`, then install the local chat + voice models:\n```bash\nollama pull qwen2.5:7b-instruct\n# Kokoro TTS voice files (~350 MB), downloaded once:\ncurl -L -o models/kokoro-v1.0.onnx \\\n  https://github.com/thewh1teagle/kokoro-onnx/releases/download/model-files-v1.0/kokoro-v1.0.onnx\ncurl -L -o models/voices-v1.0.bin \\\n  https://github.com/thewh1teagle/kokoro-onnx/releases/download/model-files-v1.0/voices-v1.0.bin\n```\nLocal STT is mlx-whisper on Apple Silicon (auto-falls back to faster-whisper elsewhere).\n</details>\n\n---\n\n## Platform support\n\n**macOS only, for now.** The voice *brain* (LiveKit + Deepgram/Cerebras, or local\nWhisper/Kokoro/Ollama) is cross-platform, but nearly every *action* JARVIS performs\nis implemented with **macOS-native tools**:\n\n| Capability | macOS tool used |\n|---|---|\n| Spotify, browser & media control, focus mode, \"which window is focused\" | **AppleScript** (`osascript`) |\n| Screen vision (screenshots) | **`screencapture`** |\n| Opening files & apps | **`open`** |\n| Alarm sound | **`afplay`** |\n| Fastest local STT | **mlx-whisper** (Apple Silicon / Metal) |\n\nSo on **Windows or Linux it will not work out of the box** — you'd get the voice\npipeline but none of the actions. A port is feasible but non-trivial: swap the macOS\nshells for platform equivalents (e.g. `spotipy`/media keys for music, Playwright/CDP\nfor the browser, `mss`/`pyautogui` for screenshots, `xdg-open`/`start` for opening\nfiles, `playsound` for alarms) and use `faster-whisper` (already the fallback)\ninstead of mlx-whisper. The cross-platform pieces (Ollama chat, RAG embeddings,\nTavily search) run anywhere, but there's no supported non-macOS entrypoint today.\nPRs welcome.\n\n---\n\n## Configuration reference\n\nAll settings live in **`.env`** (copied from **`.env.example`**, which documents\nevery option with inline notes). The essentials:\n\n| Variable | Default | Meaning |\n|---|---|---|\n| `JARVIS_MODE` | `1` | `1` cloud, `0` local |\n| `JARVIS_ORCHESTRATOR` | `openai` | full-featured single-agent brain (or `langgraph` / `native`) |\n| `JARVIS_AGENT_PROVIDER` | `cerebras` | `cerebras` (fast) or `openai`; falls back to OpenAI if no Cerebras key |\n| `JARVIS_WAKE` / `JARVIS_WAKE_WORDS` | `1` / `jarvis` | strict wake phrase on/off + trigger words |\n| `JARVIS_UI` | `1` | HUD dashboard (auto-opens on launch) |\n| `JARVIS_RAG` / `JARVIS_RAG_DIRS` | `1` / Downloads,Documents,Desktop | document indexing + folders |\n| `JARVIS_FILES` / `JARVIS_FILES_SANDBOX` | `1` / `~` | file organiser + sandbox root |\n| `JARVIS_SCHEDULER` | `1` | calendar / to-dos / reminders / alarms |\n| `JARVIS_FOCUS` / `JARVIS_FOCUS_TECHNIQUE` | `1` / `pomodoro` | focus mode + default technique |\n| `JARVIS_READING_BRIGHTNESS` / `JARVIS_READING_VOLUME` | `0.45` / `25` | reading-mode screen brightness (0–1) + system volume (0–100) |\n| `JARVIS_READING_MUSIC` | `peaceful piano instrumental for reading` | Spotify query for reading-mode background music |\n| `JARVIS_READING_DARK_MODE` / `JARVIS_READING_NIGHT_SHIFT` | `1` / `1` | dark background + warm tone in reading mode |\n| `JARVIS_BROWSER_AGENT` | `1` | logged-in browser workflows |\n\n---\n\n## Architecture\n\n```\n                 ┌─────────── LiveKit voice pipeline ───────────┐\nmic ─► VAD ─► STT ─► [wake gate] ─► BRAIN (agent + tools + memory) ─► TTS ─► speaker\n                                        │\n        ┌───────────────────────────────┼───────────────────────────────┐\n     Spotify   Browser/Web   Screen vision   Documents(RAG)   Files\n     Calendar/Reminders     Focus mode      HUD dashboard     Memory\n```\n\n- **Wake word with smart follow-up:** say *\"hey jarvis\"* at the **start** of a\n  turn to activate (the phrase is stripped). Mentions in music or background\n  speech are ignored. After a **clarifying question** it stays awake; after a\n  normal answer it sleeps. `JARVIS_WAKE=0` replies to everything.\n- **One agent, all tools:** device actions resolve in one fast round-trip;\n  informational tools keep the loop open so multi-step requests complete and are\n  synthesised into one answer. `gpt-oss` runs at low reasoning effort for speed.\n- **Custom LiveKit LLM adapter:** the whole agent framework is the LLM node of the\n  STT→LLM→TTS pipeline; STT/TTS are unchanged.\n\n---\n\n## Privacy & safety\n\n- **File organiser can never delete** — only read / copy / move / open — and is\n  **sandboxed to your home folder** (enforced structurally).\n- **Local where it matters:** document embeddings and long-term memory are computed\n  locally (Ollama) and stored as local JSON under `~/.jarvis/`.\n- **LOCAL mode** (`JARVIS_MODE=0`) keeps the entire voice loop on-device.\n- **`.env` holds secrets and is git-ignored** — never commit it.\n- **Cloud egress, made explicit:** in cloud mode, audio goes to Deepgram and your\n  text/tool calls + screenshots go to Cerebras/OpenAI. The browser agent sends\n  browsed page content to OpenAI and acts on your logged-in sessions.\n\n---\n\n## License\n\n[MIT](LICENSE) — free to use, modify, and distribute. Contributions and forks welcome.\n\n---\n\n## Troubleshooting\n\n- **No response / no audio** — check your mic in System Settings ▸ Privacy ▸\n  Microphone; make sure Ollama is running (`brew services start ollama`).\n- **\"couldn't capture the screen\"** — grant **Screen Recording** to your terminal, then reopen it.\n- **Focus mode / music does nothing** — grant **Automation** (Chrome + System\n  Events) the first time macOS prompts; if you dismissed it, re-enable under Privacy ▸ Automation.\n- **Cerebras/Deepgram errors** — re-check the keys in `.env` (no quotes, no trailing spaces).\n- **Documents not found** — put files in `~/Downloads`, `~/Documents`, or `~/Desktop`;\n  force a rebuild with `python scripts/reindex_rag.py`.\n- **Spotify won't play** — the desktop app must be installed and logged in.\n\n### Handy scripts\n- `scripts/Start JARVIS.command` — double-click launcher.\n- `scripts/make_app.sh` — build `JARVIS.app`.\n- `scripts/setup_browser_agent.sh` — install the isolated browser-agent venv.\n- `scripts/reindex_rag.py` — force a full document re-index.\n- `python -m jarvis.spotify_auth` — authorise Spotify playlists/library.\n</content>\n","readmeFilename":"README.md"}