agent-vision-toolkit

Anionex

为纯文本模型"看图“设计更好的视觉工具箱和技能,支持多图理解,图片问答,前端UI还原、GUI 自动化等,并可选无缝接入多个主流agent,直接识别粘贴图片| A vision toolkit and skill designed for text-only llms — image Q&A, long-screenshot OCR, frontend UI restoration, and GUI automation, with optional seamless integration for Codex, Claude Code, Pi, Oh My Pi, and OpenCode

agentagent-skillsclaude-codecodexcomputer-usedeepseekdsh-pluginglmharness-engineeringmultimodalopencodetext-only-llmvisionvision-language-model

安装

dsh plugin add github:Anionex/agent-vision-toolkit

首次安装 GitHub 来源的包时,可能需要在 profile 的 pnpm-workspace.yamlallowBuilds 中允许该包的构建脚本;也可用 github:Anionex/agent-vision-toolkit#<commit-sha> 锁定版本。详见官方文档

还没有 DSH?两步开始 →

第一步:准备 Node.js 环境

DSH 依赖 Node.js(建议 20 LTS 或更新版本)。终端里运行 node -v 能显示版本号即已就绪。

已有 Node.js:直接进入第二步。

没有 Node.js:去 nodejs.org 下载 LTS 安装包(Windows/macOS 双击安装);或用包管理器:

brew install node
winget install OpenJS.NodeJS.LTS

第二步:启动 DSH

无需安装,直接运行:

npx @deepseek-ai/dsh web

浏览器打开 http://127.0.0.1:3080,在 Web UI 的插件市场里搜索 agent-vision-toolkit,或粘贴上面的安装命令。

在 GitHub 查看 分享到 X