Files
Kigi-CLI/crates/codegen/kigi-pager-pty-harness/benches/pty_bench.rs
T
ZacharyZhang-NY d6c20fc13f M0: compilable skeleton — Kigi 0.1.0 fork surgery
Hard fork of xai-org/grok-build (Apache-2.0) re-targeted as Kigi, an
unofficial Kimi Code CLI community build.

Rename & identity
- 72 xai-*/xai-grok-* crates -> kigi-* (explicit: xai-grok-pager-bin ->
  kigi-bin [binary `kigi`], xai-grok-pager -> kigi-tui; rest mechanical);
  ptyctl, ptyctl-cli, third_party/ unchanged; proto package
  xai.grok.tools.v1 -> kigi.tools.v1
- Config home ~/.kigi (KIGI_SHARE_DIR override), env prefix GROK_* ->
  KIGI_*, `kigi --version` carries the unofficial-community-build notice
- clap identity, help text, startup banner, prompt templates rebranded
  (templates re-encrypted)

Deletions (PRD removal list #5/#6/#7/#9/#10)
- voice input (xai-grok-voice) and all TUI wiring
- telemetry: Mixpanel client, external OTel stream, Sentry, OTLP layers,
  trace/GCS/S3 upload queues (kigi-file-utils halved), workspace upload
  module & dc_log, heap-profile uploader, auth-diagnostics uploader,
  session-analytics halves of feedback; local zero-egress observability
  preserved in new kigi-log crate (unified log, --debug firehose,
  subsystem file logs, opt-in instrumentation)
- announcements (crate, remote-settings fields, TUI surfaces)
- plugin marketplace (crate, sources/browse/CTA/extensions-modal tab);
  direct plugin install/uninstall/update via kigi-agent git_install kept
- relay/gateway/assets endpoints and features (agent relay, headless
  relay transport, gateway bridge, LeaderEnvUrls); leader IPC socket now
  ~/.kigi/leader.sock + KIGI_LEADER_SOCKET, no ws-url derivation
- functional types rehomed instead of deleted: PermissionMode ->
  kigi-config-types, McpInitStrategy -> kigi-mcp, PrCreationSource ->
  session signals, TerminalDiagnostics -> kigi-pager-render, agent_id ->
  shell util

Endpoints
- kigi-env rewritten: single production KigiEndpoints {coding_api_base_url
  https://api.kimi.com/coding/v1 (KIGI_CODE_BASE_URL), oauth_host
  https://auth.kimi.com (KIGI_OAUTH_HOST), update_base_url (GitHub
  Releases API), upgrade_page_url}; GrokBuildEnvironment enum deleted

Toolchain & workspace hygiene
- Rust 1.97.0 pinned; edition 2024; full cargo update; git2 hoisted to
  workspace at 0.21 (Option->Result API migration), quick-xml 0.41
- Root Cargo.toml hand-maintained (PRD §8.1): version 0.1.0 inherited by
  all members, members sorted, unused deps pruned
- cargo-deny advisories gate (deny.toml with documented transitive
  exceptions); CI workflow (check/clippy/fmt/deny/test, macOS+Linux)
- cross-crate test seams re-gated behind `test-support` cargo feature;
  insta snapshot baselines renamed to the kigi_tui prefix
- clippy --workspace --all-targets: zero warnings; fmt clean

Fixes surfaced by the port
- updater probe/installer divergence (bin/kigi vs bin/grok symlink set)
- idle model-metadata refresh dead under KIGI_CODE_BASE_URL override
  (new is_effective_coding_endpoint_url, loopback+override aware)
- macOS symlinked-TMPDIR fixture canonicalization (foreign_sessions,
  fast-worktree); RSS measurement tests serialized via serial_test

Docs & legal (Apache §4)
- NOTICE added (upstream attribution + change statement); THIRD-PARTY
  notices sustained; kigi-tools ported-code notices extended; README,
  CONTRIBUTING, SECURITY, AGENTS.md rewritten

Out of scope for M0 (tracked): Kimi auth/inference (M1), search/fetch,
command parity, config import (M2), Computer Hub excision & final
brand-token sweep (M2), distribution & self-update rewrite (M3).
2026-07-17 05:31:01 -04:00

201 lines
6.0 KiB
Rust

//! `pty-bench` — PTY benchmark CLI for `kigi-tui`.
//!
//! Spawns the real pager binary in a PTY, dispatches named scenarios, and
//! emits aggregated results as JSON. Supports baseline comparison for CI
//! regression detection.
//!
//! ## Typical use
//!
//! Run a single scenario locally:
//! ```bash
//! cargo bench -p kigi-pager-pty-harness \
//! --bench pty_bench -- --scenario scroll-stress
//! ```
//!
//! Run every scenario and write a new baseline:
//! ```bash
//! cargo bench -p kigi-pager-pty-harness \
//! --bench pty_bench -- --all \
//! --write-baseline benches/pty_baselines/local.json
//! ```
//!
//! Run every scenario in CI and fail on >15% p99 regression:
//! ```bash
//! PAGER_BINARY=./artifacts/grok-${VERSION}-linux-x86_64 \
//! cargo bench -p kigi-pager-pty-harness \
//! --bench pty_bench -- --all \
//! --baseline benches/pty_baselines/linux-x86_64.json
//! ```
use std::path::PathBuf;
use std::process::ExitCode;
use anyhow::{Context, Result, bail};
use clap::Parser as ClapParser;
use kigi_pager_pty_harness::{
BenchResults, ContentController, PtyHarness, Scenario, compare_baseline, pager_binary,
results::{DEFAULT_REGRESSION_THRESHOLD, load_baseline, write_baseline},
};
#[derive(ClapParser, Debug)]
#[command(
name = "pty-bench",
about = "PTY benchmark harness for kigi-tui",
long_about = None,
)]
struct Cli {
/// Run a single scenario by name. Mutually exclusive with --all.
#[arg(long, value_enum, conflicts_with = "all")]
scenario: Option<Scenario>,
/// Run every scenario.
#[arg(long)]
all: bool,
/// Path to the pager binary. Defaults to auto-resolve (PAGER_BINARY env
/// or a locally-built debug binary).
#[arg(long)]
binary: Option<PathBuf>,
/// Terminal rows.
#[arg(long, default_value_t = 50)]
rows: u16,
/// Terminal columns.
#[arg(long, default_value_t = 120)]
cols: u16,
/// Compare results against a baseline file and exit non-zero on
/// regression (>15% p99 delta by default).
#[arg(long, value_name = "PATH")]
baseline: Option<PathBuf>,
/// Save the current run as a new baseline.
#[arg(long, value_name = "PATH", conflicts_with = "baseline")]
write_baseline: Option<PathBuf>,
/// Regression threshold as a fraction of baseline p99 (0.15 = 15%).
#[arg(long, default_value_t = DEFAULT_REGRESSION_THRESHOLD)]
threshold: f64,
/// Accepted for `cargo bench` compatibility (libtest-style argument).
/// We ignore it — this isn't a libtest harness.
#[arg(long, hide = true)]
#[allow(dead_code)]
bench: bool,
}
#[tokio::main(flavor = "multi_thread", worker_threads = 2)]
async fn main() -> ExitCode {
tracing_subscriber::fmt()
.with_writer(std::io::stderr)
.init();
match run().await {
Ok(code) => code,
Err(e) => {
eprintln!("pty-bench failed: {e:#}");
ExitCode::from(2)
}
}
}
async fn run() -> Result<ExitCode> {
let cli = Cli::parse();
let binary = match cli.binary {
Some(b) => b,
None => pager_binary().context("resolve pager binary")?,
};
let scenarios: Vec<Scenario> = if cli.all {
Scenario::ALL.to_vec()
} else if let Some(s) = cli.scenario {
vec![s]
} else {
bail!("specify --scenario <name> or --all");
};
tracing::info!(
binary = %binary.display(),
rows = cli.rows,
cols = cli.cols,
count = scenarios.len(),
"starting pty-bench run"
);
let mut results: Vec<BenchResults> = Vec::with_capacity(scenarios.len());
for scenario in scenarios {
tracing::info!(scenario = scenario.as_str(), "running scenario");
let content = ContentController::start()
.await
.context("start ContentController")?;
let mut harness =
PtyHarness::spawn_with_content(&binary, cli.rows, cli.cols, &content, &[])
.context("spawn pager PTY harness")?;
let res = scenario.run(&mut harness, &content).await;
// Best-effort cleanup regardless of scenario outcome.
let _ = harness.quit();
match res {
Ok(r) => {
tracing::info!(
scenario = %r.scenario,
frames = r.total_frames,
p50_ms = r.p50_ms,
p99_ms = r.p99_ms,
"scenario complete"
);
results.push(r);
}
Err(e) => {
tracing::warn!(scenario = scenario.as_str(), error = %e, "scenario failed");
results.push(BenchResults::from_timings(
scenario.as_str(),
&[],
std::time::Duration::ZERO,
));
}
}
}
// Emit JSON to stdout for downstream consumption.
let json = serde_json::to_string_pretty(&results).context("serialize results")?;
println!("{json}");
if let Some(path) = cli.write_baseline {
write_baseline(&path, &results)?;
eprintln!("wrote baseline to {}", path.display());
}
if let Some(path) = cli.baseline {
let baseline = load_baseline(&path)?;
let regressions = compare_baseline(&results, &baseline, cli.threshold);
if regressions.is_empty() {
eprintln!(
"OK: no scenarios regressed beyond {:.0}% of baseline p99",
cli.threshold * 100.0
);
} else {
eprintln!(
"REGRESSION: {} scenario(s) exceeded {:.0}% p99 threshold:",
regressions.len(),
cli.threshold * 100.0
);
for r in &regressions {
eprintln!(
" {}: {:.2}ms -> {:.2}ms ({:+.1}%)",
r.scenario,
r.baseline_p99_ms,
r.current_p99_ms,
r.pct_delta * 100.0
);
}
return Ok(ExitCode::from(1));
}
}
Ok(ExitCode::SUCCESS)
}