{
  "markdown": "# Kitaru Agent Skills\n\nThis repository contains agent skills for experiencing, connecting, and using\n[Kitaru](https://kitaru.ai). They support a value-first tour of the public\nreturns agent example, custom adapter and importer development, evidence-led\ninvestigations, durable annotations, versioned cohorts, evaluator selection,\nevaluator validation, and bounded replay experiments.\n\nKitaru records agent runs as evidence-rich sessions. Model and tool activity is\nrecorded when the integration exposes it, and the skills make observability gaps\nexplicit. They help a coding agent record an unsupported framework, import an\nunsupported trace format, and organize the resulting sessions into a bounded\ninvestigation without replacing human judgment.\n\nWant to see the skills in action before installing them? [Watch the 26-minute\nKitaru guided tour](https://youtu.be/aYLfzXEr2Rk). Alex starts with the Kitaru\nQuickstart, then uses `kitaru-guided-tour` on the PydanticAI returns agent example\nto review recorded sessions, define an evaluator and cohort, replay one\nimprovement, and compare the result.\n\n<p align=\"center\">\n  <a href=\"https://youtu.be/aYLfzXEr2Rk\"><img src=\"assets/kitaru-guided-tour.webp\" alt=\"Watch the Kitaru guided tour on YouTube\"></a>\n</p>\n\n## Skills\n\n| Skill | Purpose |\n|---|---|\n| [`kitaru-hosted-onboarding-tour`](skills/kitaru-hosted-onboarding-tour/SKILL.md) | Guide the controlled ZenML Pro onboarding runner through a concise, resume-safe tour that reuses exact durable state and handles pre-existing agent names without overwriting them. |\n| [`kitaru-guided-tour`](skills/kitaru-guided-tour/SKILL.md) | Give a first-time user a prepared three-session frontend review of the PydanticAI returns agent example, collect human verdicts, turn one accepted finding into a deterministic evaluator, and finish with one approved bounded replay experiment. |\n| [`kitaru-investigation`](skills/kitaru-investigation/SKILL.md) | Act as Kitaru's front door: verify setup, record sessions or import files or provider API traces, follow up on insight cards, guide human review, define one accepted behavior and cohort, select an evaluator, and offer a bounded replay experiment. |\n| [`kitaru-validate-evaluator`](skills/kitaru-validate-evaluator/SKILL.md) | Check a judge against criterion-specific human verdicts, inspect disagreements with visual summaries, and assess untouched cases without confusing abstentions or errors with quality. |\n| [`kitaru-replay-experiment`](skills/kitaru-replay-experiment/SKILL.md) | Safely test one candidate against an exact cohort and evaluator set, supervise the run, and report improved, regressed, trade-off, or inconclusive evidence without making the deployment decision. |\n| [`kitaru-importer-builder`](skills/kitaru-importer-builder/SKILL.md) | Build and locally validate a private or packaged importer for an unsupported provider or export format, with optional API fetching, conservative session joining, explicit fidelity reporting, and separately approved remote registration and smoke import. |\n| [`kitaru-adapter-builder`](skills/kitaru-adapter-builder/SKILL.md) | Select a supported provider-backed adapter or build a project-local Python or TypeScript adapter for an unsupported agent framework, with explicit recording and replay boundaries, partial-trace handling, side-effect controls, and separately approved upstream contribution. |\n\nThe workflow keeps human observations separate from agent suggestions. It uses\nthe Kitaru frontend for human review and consumes the product-owned review link\nreturned by structured investigation creation. If no returned or documented\ncompatibility URL works, it preserves the investigation and reports the broken\nproduct handoff rather than recreating the review UI in chat.\n\n## Example prompts\n\n- \"Resume the hosted Kitaru onboarding tour from what already exists in this workspace.\"\n- \"I do not have an agent yet. Show me why Kitaru is useful.\"\n- \"Give me a guided tour of Kitaru with the public returns agent example.\"\n- \"Import last week's Langfuse traces using my existing connection and generate insights.\"\n- \"Follow up on this Kitaru insight and check its supporting sessions.\"\n- \"Investigate why this Kitaru session gave a bad support answer.\"\n- \"I am new to Kitaru. Help me review one run before we investigate more.\"\n- \"Help me discover recurring failure modes in last week's agent sessions.\"\n- \"Resume investigation `INVESTIGATION_ID` and show me what remains.\"\n- \"Turn this accepted behavior and cohort into a narrow evaluator.\"\n- \"Help me validate this TypeSafe judge against my judgments, showing disagreements before the statistics.\"\n- \"Check whether this evaluator is reliable enough for our regression checks.\"\n- \"Replay this cohort with the new prompt and tell me whether it helped.\"\n- \"Run a safe experiment with history-backed tools and no live passthrough.\"\n- \"Build a private Kitaru importer for this provider's JSONL trace export.\"\n- \"These traces store each conversation turn separately. Join them safely when importing into Kitaru.\"\n- \"Help me understand whether this partial import is safe to retry.\"\n- \"Build a Kitaru adapter for this Python agent framework.\"\n- \"Add Kitaru recording and safe replay to this TypeScript agent without changing its public API.\"\n- \"Can this framework support a Kitaru adapter, including streaming and tool replay?\"\n\n## Requirements\n\nThese skills track the Kitaru CLI, MCP, SDK, and adapter contracts developed on\n[`kitaru/develop`](https://github.com/zenml-io/kitaru/tree/develop). They require\nKitaru 0.22 or newer:\n\n```bash\nuv add \"kitaru[cli,mcp,worker]>=0.22\"\n```\n\nProvider API imports, provider connections, post-import analyzers, and insight\nhandoffs require Kitaru 0.26 or newer. Existing file-based tour and investigation\npaths remain available on their earlier supported versions.\n\nEvaluator validation uses existing investigation verdicts and exact evaluation runs. TypeSafe is optional and requires the `kitaru-typesafe-evaluator` package on the worker; the skill checks its installed schema and setup. Human agreement validation does not certify the model's probabilities or deploy an automatic gate.\n\nEach skill verifies the installed version and public schema before it acts, and\nstops when the required contract is unavailable.\n\nWhen a first-time user wants to experience Kitaru before bringing an agent or\nlearning the full method, `kitaru-guided-tour` uses the PydanticAI returns agent example to\nprepare three evidence-anchored agent observations, open a frontend review for\nhuman verdicts, turn one accepted finding into a deterministic evaluator, and\nfinish with one approved bounded replay experiment. The tour pauses at the\ninvestigation review, cohort and evaluator results, and experiment result so the\nuser can inspect each durable stage. The guided starter lives in\n[`zenml-io/kitaru` PydanticAI returns agent](https://github.com/zenml-io/kitaru/tree/main/examples/python/pydantic_ai_ticket_resolver); the\ntour recognizes renamed clones and forks from stable root contents, compares\nthem with the current trusted source, resumes existing setup, and uses the\nchecked-in Langfuse JSONL without live credentials, trace regeneration, or a\npaid model call for the recorded-evidence and evaluator stages. The final\nexperiment requires a brief explanation and separate approval before model or\nlive tool execution. The coding agent prepares observations; the human supplies\nthe whole-session verdicts. Customized templates and real user evidence route\nto `kitaru-investigation`, which supports one-run review, specific-behavior\ndebugging, and bounded recurring-problem discovery. If an\nunsupported provider prevents sessions from entering Kitaru, the\n`kitaru-importer-builder` skill creates and validates the missing integration,\nthen hands usable sessions back. If a Python agent already reports to Langfuse,\nBraintrust, LangSmith, Logfire, or Arize Phoenix, `kitaru-adapter-builder`\nchecks the provider importer package's `[adapter]` extra before proposing custom\ncode. Those importer-backed adapters require Kitaru 0.24 or newer and record\nthe provider trace after the run; they cannot apply replay overrides or\nnon-passthrough tool policies. If no supported framework integration can record\nthe agent, the skill verifies the installed SDK and framework hooks before\nbuilding a project-local adapter.\n\nThe front-door journey follows Kitaru's five-step method: Observe, Judge,\nDefine, Replay, and Compare. After investigation accepts a behavior and cohort,\nit checks the installed evaluator catalog before proposing custom evaluator\ncode. For a model judge, optionally use `kitaru-validate-evaluator` to compare its decisions with human labels before relying on it. It then offers to continue with the `kitaru-replay-experiment` skill,\ncarrying the exact accepted evidence forward without making the user copy IDs.\n\nMCP is preferred, not required. For the most direct agent experience, configure\nthe native Kitaru MCP server in `standard` mode so the host can read\ninvestigations and create review, cohort, evaluator, and workflow state.\nRead-only mode supports orientation. CLI-only operation remains supported, and\nthe skill uses the structured `kitaru` CLI for local files, built-in waiting, or\noperations MCP does not expose.\n\nReview the official\n[Kitaru MCP server guide](https://docs.zenml.io/kitaru/agent-native/mcp-server)\nbefore enabling write or destructive capabilities.\n\n## Installation and usage\n\nInstall the skills with the cross-host Agent Skills installer. This is the\nrecommended route for Codex, Cursor, Claude Code, and other compatible hosts:\n\n```bash\nnpx skills add zenml-io/kitaru-skills\n```\n\nWhen the installer asks, select the host you actually use. Include\n`kitaru-guided-tour` for the public demo and `kitaru-investigation` for your own\nagents and traces. Your agent can select a skill from context, or you can ask\nfor it by name. Exact invocation syntax varies by host.\n\n`kitaru-hosted-onboarding-tour` is intended for the controlled ZenML Pro\nonboarding runner, whose image installs all skills and supplies the prepared\ntemplate and workspace connection.\n\nConfigure Kitaru MCP separately according to your host and the official Kitaru\nMCP server guide. Restart or reload the coding-agent host process or IDE after\nadding or changing MCP configuration. An already-open task cannot discover the\nnew server. Then resume from the skill's checkpoint. A missing MCP server does\nnot block a path the CLI fully supports.\n\n### Optional Claude Code plugin installation\n\nClaude Code users who prefer its plugin marketplace can install the same skills\nas a plugin:\n\n```bash\n/plugin marketplace add zenml-io/kitaru-skills\n/plugin install kitaru@kitaru\n```\n\nYou can then invoke\n`/kitaru-hosted-onboarding-tour`, `/kitaru-guided-tour`,\n`/kitaru-investigation`, `/kitaru-validate-evaluator`, `/kitaru-replay-experiment`,\n`/kitaru-importer-builder`, or `/kitaru-adapter-builder` explicitly.\n\n### Manual installation\n\nIf your host does not support the installer, copy the relevant skill directory\ninto its skills location or load the skill's `SKILL.md` as explicit project\ncontext.\n\n## Links\n\n- [Kitaru](https://kitaru.ai)\n- [Kitaru documentation](https://docs.zenml.io/kitaru)\n- [Kitaru SDK repository](https://github.com/zenml-io/kitaru)\n\n## License\n\nApache 2.0. See [LICENSE](LICENSE).\n",
  "bytes": 11379,
  "sha": "0eaa91ca59c64a5d43b1f7c73e9109eeacf8f47810eb177c9a4304d2d2b8191f",
  "repo_slug": "zenml-io/kitaru-skills",
  "fonte": "repo",
  "truncated": false,
  "api": "https://api.agentalog.com/api/listings/skl_zenml_io_kitaru_skills_kitaru_investigat_9eb8ed11/readme"
}