Flash-MoE lets you run a massive AI model called Qwen3.5-397B right on your MacBook. This model has 397 billion parameters but uses a clever mixture of experts design. It only loads the parts it needs from your SSD as you chat. The engine is built with pure C and Metal code. No Python or heavy frameworks are required. You get about 4.4 tokens per second on a MacBook Pro with 48GB of memory. This makes huge AI possible on a laptop you already own.
Anyone with an Apple Silicon Mac and enough free SSD space can run this skill. It is perfect for developers who want to test or use a very large language model locally. The skill handles all the tricky parts like streaming expert weights and running Metal shaders. You just download the model weights and run a few simple commands.
Global
mkdir -p ~/.claude/skills/flash-moe-inferenceProject
mkdir -p .claude/skills/flash-moe-inferenceSource Repository
Find Skillsvercel-labs/skills
Find and install the perfect skill to extend your AI agent
Microsoft Foundrymicrosoft/azure-skills
Build, deploy, and improve AI agents on Microsoft Foundry from start to finish
Azure Aimicrosoft/azure-skills
Search, transcribe, and analyze with Azure AI tools for smarter apps
Azure Hosted Copilot Sdkmicrosoft/azure-skills
Build, deploy, and manage your Copilot SDK apps on Azure with ease
Skill Creatoranthropics/skills
Create, test, and improve AI agent skills with easy step-by-step guidance
Image Editagentspace-so/runcomfy-agent-skills
Smart router picks the best AI model for your image editing needs
Agentspaceagentspace-so/skills
See your AI agent's live folder from any browser instantly
Openclaw Secure Linux Cloudxixu-me/skills
Secure self-hosting of OpenClaw with private control and SSH tunneling