Production-grade skills, custom MCP servers, and real projects.
Already shipped: agent-failure-analyzer (PyPI + full observability platform)
HEADER IMAGE • GROK BUILD ARSENAL
Most demos are toys. This arsenal ships real, production-minded projects + the exact reusable skills and MCPs that made them possible — all created with Grok Build 0.1 itself.
Full self-hosted dashboard + CLI + PyPI package (agent-failure-analyzer) for analyzing AI agent sessions across Grok Build, Claude Code, Cursor, LangChain, CrewAI and more. The real shipped thing.
Interactive runner for parallel subagents with scoring, synthesis, and cost tracking.
Polished control plane for discovering, installing, validating, and monitoring MCP servers + skills.
Test selection, impact analysis, flakiness detection, and coverage-guided editing at scale.
Visible reasoning agent framework for Nemotron 3 Ultra. Streams thinking + tool deltas in color, persists rich replayable JSON traces, beautiful standalone landing page + CLI.
Secure agent runtime: gateway + explainable policy + sandbox + durable skills (Hermes). Model proposes; runtime disposes. Live stdlib control plane.
Gradio playground + screenshots exercising reasoning modes, budget control, and streaming tool calling on the live Nemotron 3 Ultra endpoint.
The skills and MCP patterns here have already shipped real, depended-on tools. The core discipline (Plan Mode + arenas + skills + verification) also ports well to Claude Code, Cursor, Aider and custom harnesses.
Production CLI, web UI, 6+ framework parsers, living failure taxonomy (30+ subcats), cost analysis, Grafana/OTLP, Docker. The flagship deliverable built with this arsenal's exact patterns.
showcase/agent-observability-dashboard/showcase/nemotron-*/ and showcase/nemoclaw-runtime/Copy these into any .grok/skills/ folder. They are narrowly scoped with excellent auto-invocation descriptions.
Narrowly-focused, agent-optimized MCP servers. The highest-leverage additions for grok-build-0.1.
Ingest, classify, fingerprint, and visualize agent sessions. Full failure taxonomy + reports.
Rich architecture and dependency graphs. Identifies boundaries and hotspots.
Smart test selection, flakiness detection, impact analysis, and coverage guidance.
Visual + accessibility QA during UI work. Self-validation for the skill/MCP ecosystem.
This is a faithful client-side simulation of the exact pattern in skills/subagent-arena/SKILL.md. Click to experience parallel charters → synthesis.