DeepSeek-OCR is a smart tool that reads text from pictures, PDFs, and documents. It uses a special kind of AI that understands both images and words. This makes it great for turning scanned documents into editable text.
The tool can handle many images at once. It outputs clean markdown or plain text. You can also choose different modes for OCR, figure parsing, or grounding.
You run it on your own computer using tools like vLLM or HuggingFace. It works best with CUDA and PyTorch installed.
Global
mkdir -p ~/.claude/skills/deepseek-ocrProject
mkdir -p .claude/skills/deepseek-ocrSource Repository
Find Skillsvercel-labs/skills
Find and install the perfect skill to extend your AI agent
Microsoft Foundrymicrosoft/azure-skills
Build, deploy, and improve AI agents on Microsoft Foundry from start to finish
Azure Aimicrosoft/azure-skills
Search, transcribe, and analyze with Azure AI tools for smarter apps
Azure Hosted Copilot Sdkmicrosoft/azure-skills
Build, deploy, and manage your Copilot SDK apps on Azure with ease
Skill Creatoranthropics/skills
Create, test, and improve AI agent skills with easy step-by-step guidance
Image Editagentspace-so/runcomfy-agent-skills
Smart router picks the best AI model for your image editing needs
Agentspaceagentspace-so/skills
See your AI agent's live folder from any browser instantly
Openclaw Secure Linux Cloudxixu-me/skills
Secure self-hosting of OpenClaw with private control and SSH tunneling