Flash-MoE lets you run a massive AI model called Qwen3.5-397B right on your MacBook. This model has 397 billion parameters but uses a clever mixture of experts design. It only loads the parts it needs from your SSD as you chat. The engine is built with pure C and Metal code. No Python or heavy frameworks are required. You get about 4.4 tokens per second on a MacBook Pro with 48GB of memory. This makes huge AI possible on a laptop you already own.
Anyone with an Apple Silicon Mac and enough free SSD space can run this skill. It is perfect for developers who want to test or use a very large language model locally. The skill handles all the tricky parts like streaming expert weights and running Metal shaders. You just download the model weights and run a few simple commands.
Global
mkdir -p ~/.claude/skills/flash-moe-inferenceProject
mkdir -p .claude/skills/flash-moe-inferenceSource Repository
Find Skillsvercel-labs/skills
Find and install the perfect skill to extend your AI agent
Microsoft Foundrymicrosoft/azure-skills
Build, deploy, and improve AI agents on Microsoft Foundry from start to finish
Azure Aimicrosoft/azure-skills
Search, transcribe, and analyze with Azure AI tools for smarter apps
Azure Hosted Copilot Sdkmicrosoft/azure-skills
Build, deploy, and manage your Copilot SDK apps on Azure with ease
Triagemattpocock/skills
Triage issues with a state machine driven by clear roles and agent briefs
Handoffmattpocock/skills
Hand off your work to another AI agent with a clear summary
Image Editagentspace-so/runcomfy-agent-skills
Smart router picks the best AI model for your image editing needs
Agentspaceagentspace-so/skills
See your AI agent's live folder from any browser instantly