Hard fork of xai-org/grok-build (Apache-2.0) re-targeted as Kigi, an
unofficial Kimi Code CLI community build.
Rename & identity
- 72 xai-*/xai-grok-* crates -> kigi-* (explicit: xai-grok-pager-bin ->
kigi-bin [binary `kigi`], xai-grok-pager -> kigi-tui; rest mechanical);
ptyctl, ptyctl-cli, third_party/ unchanged; proto package
xai.grok.tools.v1 -> kigi.tools.v1
- Config home ~/.kigi (KIGI_SHARE_DIR override), env prefix GROK_* ->
KIGI_*, `kigi --version` carries the unofficial-community-build notice
- clap identity, help text, startup banner, prompt templates rebranded
(templates re-encrypted)
Deletions (PRD removal list #5/#6/#7/#9/#10)
- voice input (xai-grok-voice) and all TUI wiring
- telemetry: Mixpanel client, external OTel stream, Sentry, OTLP layers,
trace/GCS/S3 upload queues (kigi-file-utils halved), workspace upload
module & dc_log, heap-profile uploader, auth-diagnostics uploader,
session-analytics halves of feedback; local zero-egress observability
preserved in new kigi-log crate (unified log, --debug firehose,
subsystem file logs, opt-in instrumentation)
- announcements (crate, remote-settings fields, TUI surfaces)
- plugin marketplace (crate, sources/browse/CTA/extensions-modal tab);
direct plugin install/uninstall/update via kigi-agent git_install kept
- relay/gateway/assets endpoints and features (agent relay, headless
relay transport, gateway bridge, LeaderEnvUrls); leader IPC socket now
~/.kigi/leader.sock + KIGI_LEADER_SOCKET, no ws-url derivation
- functional types rehomed instead of deleted: PermissionMode ->
kigi-config-types, McpInitStrategy -> kigi-mcp, PrCreationSource ->
session signals, TerminalDiagnostics -> kigi-pager-render, agent_id ->
shell util
Endpoints
- kigi-env rewritten: single production KigiEndpoints {coding_api_base_url
https://api.kimi.com/coding/v1 (KIGI_CODE_BASE_URL), oauth_host
https://auth.kimi.com (KIGI_OAUTH_HOST), update_base_url (GitHub
Releases API), upgrade_page_url}; GrokBuildEnvironment enum deleted
Toolchain & workspace hygiene
- Rust 1.97.0 pinned; edition 2024; full cargo update; git2 hoisted to
workspace at 0.21 (Option->Result API migration), quick-xml 0.41
- Root Cargo.toml hand-maintained (PRD §8.1): version 0.1.0 inherited by
all members, members sorted, unused deps pruned
- cargo-deny advisories gate (deny.toml with documented transitive
exceptions); CI workflow (check/clippy/fmt/deny/test, macOS+Linux)
- cross-crate test seams re-gated behind `test-support` cargo feature;
insta snapshot baselines renamed to the kigi_tui prefix
- clippy --workspace --all-targets: zero warnings; fmt clean
Fixes surfaced by the port
- updater probe/installer divergence (bin/kigi vs bin/grok symlink set)
- idle model-metadata refresh dead under KIGI_CODE_BASE_URL override
(new is_effective_coding_endpoint_url, loopback+override aware)
- macOS symlinked-TMPDIR fixture canonicalization (foreign_sessions,
fast-worktree); RSS measurement tests serialized via serial_test
Docs & legal (Apache §4)
- NOTICE added (upstream attribution + change statement); THIRD-PARTY
notices sustained; kigi-tools ported-code notices extended; README,
CONTRIBUTING, SECURITY, AGENTS.md rewritten
Out of scope for M0 (tracked): Kimi auth/inference (M1), search/fetch,
command parity, config import (M2), Computer Hub excision & final
brand-token sweep (M2), distribution & self-update rewrite (M3).
308 lines
12 KiB
Rust
308 lines
12 KiB
Rust
//! Kill-handle for a session's child-process trees (`Send + Sync`).
|
|
//!
|
|
//! If a session's worker thread wedges, its `Drop`s never run and child
|
|
//! processes leak. A `ProcessScope` lives outside the worker so a supervisor
|
|
//! can `kill_all()` one session's children without restarting the host.
|
|
//!
|
|
//! # Weak-keyed registry (PID-reuse safety)
|
|
//!
|
|
//! Each child enrolls as a [`ProcessGroup`]. The scope holds only `Weak`
|
|
//! references; the spawn-site owner holds the strong `Arc`. On `kill_all`,
|
|
//! only groups whose owner is still alive (i.e. un-reaped children) upgrade
|
|
//! successfully — reaped children upgrade to `None` and are skipped. This
|
|
//! prevents `killpg` from hitting a reused PID.
|
|
//!
|
|
//! # Residual
|
|
//!
|
|
//! `kill_all` SIGKILLs but doesn't `wait` (would race the live owner).
|
|
//! A wedged owner's killed leader stays as a zombie until the host exits —
|
|
//! bounded by wedge events, not sessions.
|
|
|
|
use std::io;
|
|
use std::sync::atomic::{AtomicBool, Ordering};
|
|
use std::sync::{Arc, Mutex, OnceLock, PoisonError, Weak};
|
|
|
|
use crate::{ProcessGroup, new_process_group};
|
|
|
|
/// A `Send + Sync` kill-handle for one unit's child-process trees. Cheap to
|
|
/// clone (shares one inner via `Arc`). See the [module docs](self) for the
|
|
/// ownership model and PID-reuse safety argument.
|
|
#[derive(Clone)]
|
|
pub struct ProcessScope {
|
|
inner: Arc<ScopeInner>,
|
|
}
|
|
|
|
struct ScopeInner {
|
|
/// One `Weak` per enrolled child tree. The strong `Arc` lives at the spawn
|
|
/// site for as long as that child is alive/owned; a dead `Weak` means the
|
|
/// owner already reaped+dropped it, so killing it is neither needed nor safe.
|
|
groups: Mutex<Vec<Weak<ProcessGroup>>>,
|
|
/// Latched once [`kill_all`](ProcessScope::kill_all) has run. Nothing calls
|
|
/// `kill_all` twice (the supervisor drops its handle after), so a child
|
|
/// enrolled *after* close — a spawn that won the close/spawn race on a
|
|
/// wedged actor — would otherwise never be reaped.
|
|
/// [`register`](ProcessScope::register) checks this under the `groups` lock
|
|
/// and kills such a child on the spot instead of enrolling a leaked `Weak`.
|
|
closed: AtomicBool,
|
|
}
|
|
|
|
impl ProcessScope {
|
|
/// Create an empty scope. Infallible — group/job handles are created lazily
|
|
/// as children are enrolled.
|
|
pub fn new() -> Self {
|
|
Self {
|
|
inner: Arc::new(ScopeInner {
|
|
groups: Mutex::new(Vec::new()),
|
|
closed: AtomicBool::new(false),
|
|
}),
|
|
}
|
|
}
|
|
|
|
/// Configure `cmd` so its spawned child becomes the leader of a new process
|
|
/// group / job. Call this **before** `cmd.spawn()`, then [`enroll`] the
|
|
/// resulting child (or build the group yourself and [`register`] it).
|
|
///
|
|
/// [`enroll`]: Self::enroll
|
|
/// [`register`]: Self::register
|
|
pub fn prepare(&self, cmd: &mut tokio::process::Command) {
|
|
new_process_group(cmd);
|
|
}
|
|
|
|
/// Register an already-built process group so the scope can reap it if its
|
|
/// owner wedges. The scope keeps only a [`Weak`]; **the caller MUST keep the
|
|
/// `Arc` alive** for as long as the child is its responsibility. Dropping the
|
|
/// last `Arc` (clean reap) makes this registration a silent no-op, which is
|
|
/// what keeps [`kill_all`] PID-reuse-safe.
|
|
///
|
|
/// For spawn sites (e.g. the MCP child handle, the bash terminal) that
|
|
/// already build their own [`ProcessGroup`] and own its lifecycle; sites that
|
|
/// don't can use [`enroll`] instead.
|
|
///
|
|
/// [`kill_all`]: Self::kill_all
|
|
/// [`enroll`]: Self::enroll
|
|
pub fn register(&self, group: &Arc<ProcessGroup>) {
|
|
let mut groups = self.lock();
|
|
if self.inner.closed.load(Ordering::Relaxed) {
|
|
// The scope was already reclaimed (`kill_all` ran) and won't run
|
|
// again, so reap this just-spawned child now rather than enroll a
|
|
// `Weak` that would leak. Closes the close/spawn race where a
|
|
// (possibly wedged) actor's spawn lands after teardown. `killpg` is
|
|
// non-blocking, so killing under the lock is fine and serializes
|
|
// with a concurrent `kill_all`.
|
|
let _ = group.kill();
|
|
return;
|
|
}
|
|
groups.retain(|w| w.strong_count() > 0);
|
|
groups.push(Arc::downgrade(group));
|
|
}
|
|
|
|
/// Build a process group for an already-spawned `child` (which must have been
|
|
/// configured via [`prepare`]), register a [`Weak`] into this scope, and
|
|
/// return the owning `Arc` for the caller to hold. Errors only if the child
|
|
/// already exited (nothing to attach) — not a leak.
|
|
///
|
|
/// [`prepare`]: Self::prepare
|
|
#[must_use = "the returned Arc<ProcessGroup> must be kept alive or the scope cannot reap the child"]
|
|
pub fn enroll(&self, child: &tokio::process::Child) -> io::Result<Arc<ProcessGroup>> {
|
|
let mut group = ProcessGroup::new()?;
|
|
group.attach(child)?;
|
|
let group = Arc::new(group);
|
|
self.register(&group);
|
|
Ok(group)
|
|
}
|
|
|
|
/// Convenience for simple sites: [`prepare`] + spawn + [`enroll`]. Returns
|
|
/// the child together with the owning `Arc<ProcessGroup>`, which the caller
|
|
/// must keep alive for the scope to be able to reap the child.
|
|
///
|
|
/// [`prepare`]: Self::prepare
|
|
/// [`enroll`]: Self::enroll
|
|
#[must_use = "the returned Arc<ProcessGroup> must be kept alive or the scope cannot reap the child"]
|
|
pub fn spawn(
|
|
&self,
|
|
mut cmd: tokio::process::Command,
|
|
) -> io::Result<(tokio::process::Child, Arc<ProcessGroup>)> {
|
|
self.prepare(&mut cmd);
|
|
#[allow(clippy::disallowed_methods)]
|
|
// ProcessScope::spawn is the enrollment primitive itself.
|
|
let child = cmd.spawn()?;
|
|
let group = self.enroll(&child)?;
|
|
Ok((child, group))
|
|
}
|
|
|
|
/// Idempotently kill every still-owned process tree (`killpg(SIGKILL)` /
|
|
/// `TerminateJobObject`). Safe to call multiple times and from any thread.
|
|
/// Groups whose owner already reaped+dropped them upgrade to `None` and are
|
|
/// skipped — so this never `killpg`s a reused PID.
|
|
pub fn kill_all(&self) {
|
|
let mut groups = self.lock();
|
|
for weak in groups.iter() {
|
|
if let Some(group) = weak.upgrade() {
|
|
// A failed kill means the group already exited (ESRCH) — benign,
|
|
// and there is nothing actionable to log at this primitive layer.
|
|
// Callers that care about the wedge-reclaim event log it there.
|
|
let _ = group.kill();
|
|
}
|
|
}
|
|
// Every weak has been handled above; clear the set.
|
|
groups.clear();
|
|
// Latch closed under the lock: a concurrent `register` either already
|
|
// pushed (its group was just killed in the loop) or now sees `closed`
|
|
// and kills its own child — nothing slips through after teardown.
|
|
self.inner.closed.store(true, Ordering::Relaxed);
|
|
}
|
|
|
|
/// Lock the group set, tolerating a poisoned mutex: the critical sections
|
|
/// here are panic-free, and a best-effort reaper must still run even if some
|
|
/// unrelated thread panicked while holding the lock.
|
|
fn lock(&self) -> std::sync::MutexGuard<'_, Vec<Weak<ProcessGroup>>> {
|
|
self.inner
|
|
.groups
|
|
.lock()
|
|
.unwrap_or_else(PoisonError::into_inner)
|
|
}
|
|
|
|
/// Number of still-live enrolled groups: weaks whose owning `Arc` has not
|
|
/// yet been dropped. This is exactly the set `kill_all` would `killpg`, so a
|
|
/// spawn-site owner that has reaped its child (and dropped its `Arc`) is no
|
|
/// longer counted — letting callers assert PID-reuse safety.
|
|
pub fn live_count(&self) -> usize {
|
|
self.lock().iter().filter(|w| w.strong_count() > 0).count()
|
|
}
|
|
}
|
|
|
|
impl Default for ProcessScope {
|
|
fn default() -> Self {
|
|
Self::new()
|
|
}
|
|
}
|
|
|
|
/// Process-global [`ProcessScope`] for reaping detached child trees at process
|
|
/// exit. Spawn sites (e.g. the local terminal backend) enroll their children
|
|
/// here, and the TUI exit paths call [`ProcessScope::kill_all`] on it so a
|
|
/// `setsid`-detached background command cannot outlive the process that started
|
|
/// it. `kill_all` latches the scope closed, so call it only when the process is
|
|
/// genuinely exiting.
|
|
pub fn global_process_scope() -> &'static ProcessScope {
|
|
static GLOBAL: OnceLock<ProcessScope> = OnceLock::new();
|
|
GLOBAL.get_or_init(ProcessScope::new)
|
|
}
|
|
|
|
impl Drop for ScopeInner {
|
|
fn drop(&mut self) {
|
|
// RAII backstop: if the last scope handle drops without an explicit
|
|
// `kill_all`, still reap any group whose owner is alive (a wedged
|
|
// unit). Dead weaks (clean teardown) upgrade to None and are skipped,
|
|
// so this stays PID-reuse-safe.
|
|
let groups = self.groups.lock().unwrap_or_else(PoisonError::into_inner);
|
|
for weak in groups.iter() {
|
|
if let Some(group) = weak.upgrade() {
|
|
let _ = group.kill();
|
|
}
|
|
}
|
|
}
|
|
}
|
|
|
|
#[cfg(all(test, unix))]
|
|
mod tests {
|
|
use super::*;
|
|
use std::time::Duration;
|
|
|
|
fn sleeper() -> tokio::process::Command {
|
|
let mut c = tokio::process::Command::new("sleep");
|
|
c.arg("1000");
|
|
c
|
|
}
|
|
|
|
/// `wait()` completing == the process actually died (and is reaped). If the
|
|
/// kill failed, the `sleep 1000` would run on and `wait()` would time out.
|
|
async fn died(child: &mut tokio::process::Child) -> bool {
|
|
tokio::time::timeout(Duration::from_secs(3), child.wait())
|
|
.await
|
|
.is_ok()
|
|
}
|
|
|
|
#[tokio::test]
|
|
async fn kill_all_reaps_every_enrolled_child() {
|
|
let scope = ProcessScope::new();
|
|
// The owner (here, the test) keeps the Arcs alive — as a live spawn site
|
|
// would for as long as the child is running.
|
|
let (mut c1, _g1) = scope.spawn(sleeper()).unwrap();
|
|
let (mut c2, _g2) = scope.spawn(sleeper()).unwrap();
|
|
assert_eq!(scope.live_count(), 2);
|
|
|
|
scope.kill_all();
|
|
assert!(died(&mut c1).await, "child 1 must die after kill_all");
|
|
assert!(died(&mut c2).await, "child 2 must die after kill_all");
|
|
}
|
|
|
|
#[tokio::test]
|
|
async fn kill_all_is_idempotent() {
|
|
let scope = ProcessScope::new();
|
|
let (mut c, _g) = scope.spawn(sleeper()).unwrap();
|
|
scope.kill_all();
|
|
scope.kill_all(); // second call must not panic / error
|
|
assert!(died(&mut c).await);
|
|
}
|
|
|
|
/// PID-reuse safety: once the owner drops its `Arc` (simulating a clean
|
|
/// reap), the scope no longer references the group and `kill_all` is a no-op
|
|
/// for it — it must NOT `killpg` a now-reapable/reused group id.
|
|
#[tokio::test]
|
|
async fn kill_all_skips_group_whose_owner_dropped() {
|
|
let scope = ProcessScope::new();
|
|
let (mut c, group) = scope.spawn(sleeper()).unwrap();
|
|
assert_eq!(scope.live_count(), 1);
|
|
|
|
// Owner reaps + releases ownership.
|
|
drop(group);
|
|
assert_eq!(
|
|
scope.live_count(),
|
|
0,
|
|
"dropping the owner's Arc must make the scope's weak dead"
|
|
);
|
|
|
|
scope.kill_all(); // must be a no-op for the now-unowned group
|
|
// The child was never killed by the scope; clean it up so the test
|
|
// doesn't leak a real `sleep` process.
|
|
let _ = c.start_kill();
|
|
let _ = c.wait().await;
|
|
}
|
|
|
|
#[tokio::test]
|
|
async fn drop_reaps_children_while_owner_alive() {
|
|
let scope = ProcessScope::new();
|
|
// Owner Arc is held by the test, so the scope's weak is live; dropping
|
|
// the scope's last handle must SIGKILL the still-owned group.
|
|
let (mut c, _g) = scope.spawn(sleeper()).unwrap();
|
|
drop(scope);
|
|
assert!(
|
|
died(&mut c).await,
|
|
"dropping the scope must reap a still-owned enrolled child"
|
|
);
|
|
}
|
|
|
|
/// Close/spawn race: a child enrolled *after* `kill_all` must be killed on
|
|
/// the spot, not leaked — `kill_all` won't run again to catch it.
|
|
#[tokio::test]
|
|
async fn register_after_kill_all_reaps_immediately() {
|
|
let scope = ProcessScope::new();
|
|
scope.kill_all(); // close the scope
|
|
|
|
let mut cmd = sleeper();
|
|
scope.prepare(&mut cmd);
|
|
#[allow(clippy::disallowed_methods)] // test: exercises enroll() after close
|
|
let mut child = cmd.spawn().unwrap();
|
|
let _group = scope.enroll(&child).unwrap(); // register() runs post-close
|
|
assert_eq!(
|
|
scope.live_count(),
|
|
0,
|
|
"a post-close register must not enroll the group"
|
|
);
|
|
assert!(
|
|
died(&mut child).await,
|
|
"a child registered after kill_all must be killed immediately"
|
|
);
|
|
}
|
|
}
|