M0: compilable skeleton — Kigi 0.1.0 fork surgery

Hard fork of xai-org/grok-build (Apache-2.0) re-targeted as Kigi, an
unofficial Kimi Code CLI community build.

Rename & identity
- 72 xai-*/xai-grok-* crates -> kigi-* (explicit: xai-grok-pager-bin ->
  kigi-bin [binary `kigi`], xai-grok-pager -> kigi-tui; rest mechanical);
  ptyctl, ptyctl-cli, third_party/ unchanged; proto package
  xai.grok.tools.v1 -> kigi.tools.v1
- Config home ~/.kigi (KIGI_SHARE_DIR override), env prefix GROK_* ->
  KIGI_*, `kigi --version` carries the unofficial-community-build notice
- clap identity, help text, startup banner, prompt templates rebranded
  (templates re-encrypted)

Deletions (PRD removal list #5/#6/#7/#9/#10)
- voice input (xai-grok-voice) and all TUI wiring
- telemetry: Mixpanel client, external OTel stream, Sentry, OTLP layers,
  trace/GCS/S3 upload queues (kigi-file-utils halved), workspace upload
  module & dc_log, heap-profile uploader, auth-diagnostics uploader,
  session-analytics halves of feedback; local zero-egress observability
  preserved in new kigi-log crate (unified log, --debug firehose,
  subsystem file logs, opt-in instrumentation)
- announcements (crate, remote-settings fields, TUI surfaces)
- plugin marketplace (crate, sources/browse/CTA/extensions-modal tab);
  direct plugin install/uninstall/update via kigi-agent git_install kept
- relay/gateway/assets endpoints and features (agent relay, headless
  relay transport, gateway bridge, LeaderEnvUrls); leader IPC socket now
  ~/.kigi/leader.sock + KIGI_LEADER_SOCKET, no ws-url derivation
- functional types rehomed instead of deleted: PermissionMode ->
  kigi-config-types, McpInitStrategy -> kigi-mcp, PrCreationSource ->
  session signals, TerminalDiagnostics -> kigi-pager-render, agent_id ->
  shell util

Endpoints
- kigi-env rewritten: single production KigiEndpoints {coding_api_base_url
  https://api.kimi.com/coding/v1 (KIGI_CODE_BASE_URL), oauth_host
  https://auth.kimi.com (KIGI_OAUTH_HOST), update_base_url (GitHub
  Releases API), upgrade_page_url}; GrokBuildEnvironment enum deleted

Toolchain & workspace hygiene
- Rust 1.97.0 pinned; edition 2024; full cargo update; git2 hoisted to
  workspace at 0.21 (Option->Result API migration), quick-xml 0.41
- Root Cargo.toml hand-maintained (PRD §8.1): version 0.1.0 inherited by
  all members, members sorted, unused deps pruned
- cargo-deny advisories gate (deny.toml with documented transitive
  exceptions); CI workflow (check/clippy/fmt/deny/test, macOS+Linux)
- cross-crate test seams re-gated behind `test-support` cargo feature;
  insta snapshot baselines renamed to the kigi_tui prefix
- clippy --workspace --all-targets: zero warnings; fmt clean

Fixes surfaced by the port
- updater probe/installer divergence (bin/kigi vs bin/grok symlink set)
- idle model-metadata refresh dead under KIGI_CODE_BASE_URL override
  (new is_effective_coding_endpoint_url, loopback+override aware)
- macOS symlinked-TMPDIR fixture canonicalization (foreign_sessions,
  fast-worktree); RSS measurement tests serialized via serial_test

Docs & legal (Apache §4)
- NOTICE added (upstream attribution + change statement); THIRD-PARTY
  notices sustained; kigi-tools ported-code notices extended; README,
  CONTRIBUTING, SECURITY, AGENTS.md rewritten

Out of scope for M0 (tracked): Kimi auth/inference (M1), search/fetch,
command parity, config import (M2), Computer Hub excision & final
brand-token sweep (M2), distribution & self-update rewrite (M3).
This commit is contained in:
2026-07-17 05:31:01 -04:00
commit d6c20fc13f
2612 changed files with 1353757 additions and 0 deletions
+354
View File
@@ -0,0 +1,354 @@
//! Per-`Ready`-client transport-closed poller.
//!
//! Each successful handshake spawns one [`TransportLivenessHandle`]
//! that polls the owning [`McpClient`]'s state machine on a small
//! interval (default 500 ms). On the **first observation of
//! `Ready` + `is_transport_closed() == true`** it emits a single
//! [`McpClientEvent::TransportClosed`] and exits.
//!
//! The poller is *one-shot*. The session-side dispatcher decides
//! what to do with the event — drop the dead client, surface
//! `unavailable` over ACP, and trigger a restart on a debounce.
//!
//! ## Watcher state machine
//!
//! Per-tick classification (single state-mutex acquisition via
//! [`McpClient::liveness_check`]):
//!
//! | State observed | Action | Emit? |
//! |-------------------------------|-------------------|------------------------|
//! | `Ready` + transport open | continue polling | no |
//! | `Ready` + transport closed | clear slot, exit | `TransportClosed` |
//! | `Initializing` (re-handshake) | clear slot, exit | no — silent withdrawal |
//! | `Pending` | clear slot, exit | no — silent withdrawal |
//! | `Empty` | clear slot, exit | no — silent withdrawal |
//!
//! This avoids the previous false-positive `TransportClosed` whenever
//! someone called `reset_transport()` or any other code path
//! moved the state away from `Ready`.
//!
//! ## Slot-clearing on exit
//!
//! Before exiting, the task clears
//! [`McpClient::liveness_handle`] so a subsequent
//! [`McpClient::arm_liveness_watcher`] call can install a fresh
//! handle. Without this, a dead-but-still-present
//! [`TransportLivenessHandle`] would silently block re-arming.
//!
//! ## Cancellation
//!
//! Dropping the [`TransportLivenessHandle`] cancels the spawned
//! task via [`tokio_util::sync::DropGuard`]. Both teardown paths
//! (slot-clear-from-inside, external drop) end with the same
//! handle-drop semantics.
//!
//! ## Why polling, not a `JoinHandle`-on-the-service-loop?
//!
//! rmcp 2.1's `RunningService` does not expose a future that
//! resolves on transport shutdown. The closest signal is
//! `Peer::is_transport_closed()` (a state inspection), which is the
//! same one [`McpClient::is_healthy`] reads. A `select!` on a
//! per-client `Notify` would require patching rmcp; polling avoids
//! that quarantine break and the overhead is negligible (one mutex
//! acquire + one atomic load per tick).
use std::sync::Arc;
use std::time::Duration;
use tokio::sync::mpsc::UnboundedSender;
use tokio_util::sync::{CancellationToken, DropGuard};
use crate::servers::{LivenessCheck, McpClient, McpClientEvent, McpServerName};
/// Default poll interval. Picked to keep mean detection latency
/// under one second while polling is cheap (`Mutex::lock` +
/// `tokio::sync::mpsc::is_closed`). See module doc.
pub const DEFAULT_POLL_INTERVAL: Duration = Duration::from_millis(500);
/// Shared liveness-handle slot type — same Arc lives on the
/// [`McpClient`] and is passed into the polling task so the task
/// can clear the slot before exiting. Kept private to the crate to
/// discourage external mutation.
pub(crate) type SharedLivenessSlot = Arc<parking_lot::Mutex<Option<TransportLivenessHandle>>>;
/// Release the client's liveness slot, dropping any handle it held.
///
/// Both watcher-exit arms (transport closed, transient state drift)
/// clear the slot so a later [`McpClient::arm_liveness_watcher`] can
/// install a fresh handle. The taken handle is dropped outside the
/// critical section — the lock is held for nanoseconds.
fn clear_liveness_slot(slot: &SharedLivenessSlot) {
let stale_handle = slot.lock().take();
drop(stale_handle);
}
/// RAII handle for the per-client liveness task.
///
/// Drop semantics: drop → `DropGuard` cancels the `CancellationToken`
/// → the polling task wakes from `select!` on the next tick and
/// exits cleanly without emitting. There is no public `abort()` /
/// `stop()` — the contract is "tie the handle to the client".
pub struct TransportLivenessHandle {
/// Name of the server this handle is watching. Exposed for
/// diagnostics / log lines.
pub server_name: McpServerName,
/// On drop, cancels the spawned task. Field is held purely for
/// its `Drop`; never read.
_cancel: DropGuard,
}
impl std::fmt::Debug for TransportLivenessHandle {
fn fmt(&self, f: &mut std::fmt::Formatter<'_>) -> std::fmt::Result {
f.debug_struct("TransportLivenessHandle")
.field("server_name", &self.server_name)
.finish()
}
}
impl TransportLivenessHandle {
pub fn server_name(&self) -> &str {
&self.server_name
}
}
/// Spawn a one-shot transport-liveness poller for a `Ready` client.
///
/// # Parameters
///
/// - `server_name`: bound to emitted events.
/// - `client`: `Arc<McpClient>` whose `liveness_check` we poll.
/// - `poll_interval`: tick period.
/// - `on_event`: sink for `TransportClosed` if observed.
/// - `liveness_slot`: shared Arc to the owning `McpClient`'s
/// `liveness_handle` field. Cleared from inside the task before
/// exit.
///
/// # Contract
///
/// - Caller MUST have already observed the client transition to
/// [`crate::servers::ClientStateKind::Ready`]
/// ([`McpClient::arm_liveness_watcher`] enforces this).
/// - The poller exits silently on transient non-`Ready` states; only
/// `Ready` + closed transport produces an event.
/// - The send may fail if the dispatcher has dropped its receiver
/// (subagent teardown, session shutdown). That's logged at debug
/// and the task exits — there's no retry.
///
/// # Why `tokio::time::interval` and not `sleep_until`
///
/// `interval` ticks immediately on first poll, which gives us
/// instant detection of "the transport was already closed when the
/// handle was spawned" — a real failure mode if a handshake races a
/// shutdown event from the server (e.g. Ctrl+C against an stdio
/// server that died between `Ready` write and the spawn). The
/// `MissedTickBehavior::Skip` default is fine: the worst case under
/// a runtime stall is "we don't poll for a while", which only
/// delays detection.
pub fn spawn_transport_liveness(
server_name: McpServerName,
client: Arc<McpClient>,
poll_interval: Duration,
on_event: UnboundedSender<McpClientEvent>,
liveness_slot: SharedLivenessSlot,
) -> TransportLivenessHandle {
let token = CancellationToken::new();
let drop_guard = token.clone().drop_guard();
let server_name_for_task = server_name.clone();
tokio::spawn(async move {
let mut tick = tokio::time::interval(poll_interval);
// Skip missed ticks under runtime stall — see fn doc.
tick.set_missed_tick_behavior(tokio::time::MissedTickBehavior::Skip);
loop {
tokio::select! {
_ = token.cancelled() => {
// Cancelled by the handle's `DropGuard`. The
// caller dropped the handle (e.g. McpClient
// teardown), so the slot has already been
// mutated externally — do not race the dropper
// by clearing the slot here.
tracing::trace!(
server = %server_name_for_task,
"transport liveness watcher cancelled by handle drop",
);
return;
}
_ = tick.tick() => {
match client.liveness_check().await {
LivenessCheck::Healthy => continue,
LivenessCheck::TransportClosed => {
tracing::info!(
server = %server_name_for_task,
"transport liveness watcher detected closed transport",
);
// Clear our own slot before exiting so a
// subsequent `arm_liveness_watcher` can
// install a fresh handle.
//
// Self-cancel-by-drop: clearing the slot
// drops the taken `TransportLivenessHandle`,
// whose `DropGuard` cancels the very
// `CancellationToken` this task is
// `select!`ing on. Benign because we
// `return` immediately — but DO NOT add any
// post-`return` work that re-enters the
// `select!`; it would race this self-cancel.
clear_liveness_slot(&liveness_slot);
if on_event
.send(McpClientEvent::TransportClosed {
server: server_name_for_task.clone(),
// Bind the event to THIS client
// instance so the dispatcher can
// skip evicting a replacement
// registered under the same name.
client_id: client.client_id(),
})
.is_err()
{
tracing::debug!(
server = %server_name_for_task,
"dispatcher receiver dropped; liveness watcher exiting silently",
);
}
return;
}
LivenessCheck::Transient => {
// State moved out of `Ready` (re-handshake
// started, or the transport was reset
// externally). The watcher detects
// *transport closure*, not state changes,
// so exit silently; the caller re-arms a
// fresh watcher when the new handshake
// completes.
tracing::debug!(
server = %server_name_for_task,
"transport liveness watcher: state drifted out of Ready, exiting silently",
);
clear_liveness_slot(&liveness_slot);
return;
}
}
}
}
}
});
TransportLivenessHandle {
server_name,
_cancel: drop_guard,
}
}
#[cfg(test)]
mod tests {
use super::*;
use crate::servers::McpClient;
use tokio::sync::mpsc::unbounded_channel;
/// Stub client whose `liveness_check()` returns
/// `LivenessCheck::Transient`: `McpClient::stub` lands in
/// `ClientState::Empty`, which the liveness classifier treats as a
/// silent-withdrawal state (NOT `TransportClosed`).
fn make_stub_client() -> Arc<McpClient> {
Arc::new(McpClient::stub("test-server"))
}
/// Contract: a watcher whose owning client never reaches
/// `Ready+closed` (here the stub is `Empty`) exits **silently**
/// — no `TransportClosed` event, and the slot is cleared.
/// The watcher must not false-positive on non-`Ready` states.
#[tokio::test(start_paused = true)]
async fn poller_silent_exit_on_non_ready_state() {
let (tx, mut rx) = unbounded_channel::<McpClientEvent>();
let slot: SharedLivenessSlot = Arc::new(parking_lot::Mutex::new(None));
let client = make_stub_client();
let handle = spawn_transport_liveness(
"test-server".to_string(),
client,
Duration::from_millis(500),
tx,
Arc::clone(&slot),
);
// Pre-populate the slot so we can assert the watcher
// clears it on exit.
*slot.lock() = Some(handle);
// First `interval.tick()` fires immediately under paused
// time. The watcher classifies `Empty` as `Transient` and
// exits silently.
tokio::time::advance(Duration::from_millis(10)).await;
tokio::task::yield_now().await;
// No event emitted: the watcher exited silently.
assert!(
rx.try_recv().is_err(),
"non-Ready states must not produce TransportClosed",
);
// Slot is cleared so re-arming wouldn't be blocked.
assert!(
slot.lock().is_none(),
"watcher must clear its own slot on exit",
);
}
/// Contract: when the watcher emits `TransportClosed` it both
/// (a) sends the event and (b) clears the shared liveness
/// slot so the next `arm_liveness_watcher` succeeds.
///
/// We exercise this by constructing a client that *would*
/// classify as `Ready + closed` — but `McpClient::stub` is
/// `Empty`, which classifies as `Transient`, so this test
/// instead asserts the silent-exit path. The Ready+closed
/// path is covered by the integration test in `servers.rs`.
#[tokio::test(start_paused = true)]
async fn poller_clears_slot_on_exit() {
let (tx, mut rx) = unbounded_channel::<McpClientEvent>();
let slot: SharedLivenessSlot = Arc::new(parking_lot::Mutex::new(None));
let client = make_stub_client();
let handle = spawn_transport_liveness(
"test-server".to_string(),
client,
Duration::from_millis(500),
tx,
Arc::clone(&slot),
);
*slot.lock() = Some(handle);
tokio::time::advance(Duration::from_millis(10)).await;
tokio::task::yield_now().await;
// Advancing several intervals confirms the watcher exited
// (not just stuck in a loop without progress).
tokio::time::advance(Duration::from_secs(5)).await;
tokio::task::yield_now().await;
assert!(rx.try_recv().is_err());
assert!(
slot.lock().is_none(),
"slot must be cleared even on the silent-exit path",
);
}
/// Contract: dropping the handle stops the task without
/// emitting. Drops happen via the external `DropGuard` path,
/// distinct from the in-task slot-clear path tested above.
#[tokio::test(start_paused = true)]
async fn drop_cancels_task_before_first_tick() {
let (tx, mut rx) = unbounded_channel::<McpClientEvent>();
let slot: SharedLivenessSlot = Arc::new(parking_lot::Mutex::new(None));
let client = make_stub_client();
let handle = spawn_transport_liveness(
"test-server".to_string(),
client,
Duration::from_secs(60), // Long interval so the first tick is far away.
tx,
Arc::clone(&slot),
);
// Drop before the tick can fire — the `DropGuard` arm
// wins the `select!`.
drop(handle);
tokio::task::yield_now().await;
assert!(rx.try_recv().is_err());
}
}