Files
Kigi-CLI/crates/codegen/kigi-shell/skills/best-of-n/SKILL.md
T
ZacharyZhang-NY d6c20fc13f M0: compilable skeleton — Kigi 0.1.0 fork surgery
Hard fork of xai-org/grok-build (Apache-2.0) re-targeted as Kigi, an
unofficial Kimi Code CLI community build.

Rename & identity
- 72 xai-*/xai-grok-* crates -> kigi-* (explicit: xai-grok-pager-bin ->
  kigi-bin [binary `kigi`], xai-grok-pager -> kigi-tui; rest mechanical);
  ptyctl, ptyctl-cli, third_party/ unchanged; proto package
  xai.grok.tools.v1 -> kigi.tools.v1
- Config home ~/.kigi (KIGI_SHARE_DIR override), env prefix GROK_* ->
  KIGI_*, `kigi --version` carries the unofficial-community-build notice
- clap identity, help text, startup banner, prompt templates rebranded
  (templates re-encrypted)

Deletions (PRD removal list #5/#6/#7/#9/#10)
- voice input (xai-grok-voice) and all TUI wiring
- telemetry: Mixpanel client, external OTel stream, Sentry, OTLP layers,
  trace/GCS/S3 upload queues (kigi-file-utils halved), workspace upload
  module & dc_log, heap-profile uploader, auth-diagnostics uploader,
  session-analytics halves of feedback; local zero-egress observability
  preserved in new kigi-log crate (unified log, --debug firehose,
  subsystem file logs, opt-in instrumentation)
- announcements (crate, remote-settings fields, TUI surfaces)
- plugin marketplace (crate, sources/browse/CTA/extensions-modal tab);
  direct plugin install/uninstall/update via kigi-agent git_install kept
- relay/gateway/assets endpoints and features (agent relay, headless
  relay transport, gateway bridge, LeaderEnvUrls); leader IPC socket now
  ~/.kigi/leader.sock + KIGI_LEADER_SOCKET, no ws-url derivation
- functional types rehomed instead of deleted: PermissionMode ->
  kigi-config-types, McpInitStrategy -> kigi-mcp, PrCreationSource ->
  session signals, TerminalDiagnostics -> kigi-pager-render, agent_id ->
  shell util

Endpoints
- kigi-env rewritten: single production KigiEndpoints {coding_api_base_url
  https://api.kimi.com/coding/v1 (KIGI_CODE_BASE_URL), oauth_host
  https://auth.kimi.com (KIGI_OAUTH_HOST), update_base_url (GitHub
  Releases API), upgrade_page_url}; GrokBuildEnvironment enum deleted

Toolchain & workspace hygiene
- Rust 1.97.0 pinned; edition 2024; full cargo update; git2 hoisted to
  workspace at 0.21 (Option->Result API migration), quick-xml 0.41
- Root Cargo.toml hand-maintained (PRD §8.1): version 0.1.0 inherited by
  all members, members sorted, unused deps pruned
- cargo-deny advisories gate (deny.toml with documented transitive
  exceptions); CI workflow (check/clippy/fmt/deny/test, macOS+Linux)
- cross-crate test seams re-gated behind `test-support` cargo feature;
  insta snapshot baselines renamed to the kigi_tui prefix
- clippy --workspace --all-targets: zero warnings; fmt clean

Fixes surfaced by the port
- updater probe/installer divergence (bin/kigi vs bin/grok symlink set)
- idle model-metadata refresh dead under KIGI_CODE_BASE_URL override
  (new is_effective_coding_endpoint_url, loopback+override aware)
- macOS symlinked-TMPDIR fixture canonicalization (foreign_sessions,
  fast-worktree); RSS measurement tests serialized via serial_test

Docs & legal (Apache §4)
- NOTICE added (upstream attribution + change statement); THIRD-PARTY
  notices sustained; kigi-tools ported-code notices extended; README,
  CONTRIBUTING, SECURITY, AGENTS.md rewritten

Out of scope for M0 (tracked): Kimi auth/inference (M1), search/fetch,
command parity, config import (M2), Computer Hub excision & final
brand-token sweep (M2), distribution & self-update rewrite (M3).
2026-07-17 05:31:01 -04:00

3.4 KiB

name, description, metadata
name description metadata
best-of-n Implement a task N ways in parallel and pick the best. Spawns multiple subagents in isolated worktrees, evaluates all candidates, and applies the winner. Use when asked to "best of n", "try multiple approaches", "parallel implementations", "/best-of-n", or "/bon".
short-description
Parallel implementation tournament

/best-of-n -- Parallel Implementation Tournament

Implement a task multiple different ways in parallel, evaluate all candidates, and apply the best one.

Usage

/best-of-n [N] <task>

  • If the first token is a number 2-10, it sets the candidate count; the rest is the task.
  • If omitted, N defaults to 3.

Examples:

  • /best-of-n implement the login page (3 candidates)
  • /best-of-n 5 refactor the auth module (5 candidates)

Steps

  1. Parse the user's message to extract N (candidate count, default 3) and the task description.

  2. Spawn N subagents in a single message (parallel tool calls). Use the task tool for each with:

    • subagent_type: "general-purpose"
    • isolation: "worktree"
    • run_in_background: true
    • description: "Candidate <number>"
    • prompt: the task description, plus "You are candidate <number> of <N> independent implementations. Implement the task fully. When done, summarize your approach and the changes you made."
  3. Wait for all candidates to complete using get_task_output with block: true or wait_tasks with mode: "wait_all".

  4. Evaluate and pick the winner using the criteria below.

  5. Apply the winner's changes from its worktree to the main workspace. Review the changes in context and fix any remaining issues.

  6. End your response with WINNER: <number> (1-N).

Evaluation Criteria

Evaluate each candidate on these axes, in order of importance:

  1. Correctness -- Does the candidate actually solve the task? Does it handle the requirements completely, or does it miss important aspects? Are there logic errors, type errors, or broken imports?

  2. Code Quality -- Is the code clean, readable, and well-structured? Does it follow the patterns and conventions of the surrounding codebase? Does it avoid unnecessary complexity?

  3. Safety -- Does the candidate avoid introducing bugs, security issues, or breaking changes to existing functionality?

How to Decide

  • Focus on correctness first. A candidate that fully solves the task with minor style issues beats one that is beautifully written but incomplete or wrong.
  • If multiple candidates are equally correct, prefer the one with cleaner code and better codebase integration.
  • If a candidate introduces unnecessary changes beyond the task scope, count that against it.
  • If all candidates are poor, still pick the least bad one.

Presenting Your Evaluation

Before announcing your choice, present a structured comparison:

Dimension Candidate 1 Candidate 2 ...
Correctness Short verdict Short verdict ...
Code Quality Short verdict Short verdict ...
Safety Short verdict Short verdict ...

Then list key findings:

Finding Severity Candidate 1 Candidate 2 ...
Specific issue High/Medium/Low How handled How handled ...

State which candidate you chose and why.