Files
Kigi-CLI/crates/codegen/kigi-tty-utils/src/process_scope.rs
T
ZacharyZhang-NY d6c20fc13f M0: compilable skeleton — Kigi 0.1.0 fork surgery
Hard fork of xai-org/grok-build (Apache-2.0) re-targeted as Kigi, an
unofficial Kimi Code CLI community build.

Rename & identity
- 72 xai-*/xai-grok-* crates -> kigi-* (explicit: xai-grok-pager-bin ->
  kigi-bin [binary `kigi`], xai-grok-pager -> kigi-tui; rest mechanical);
  ptyctl, ptyctl-cli, third_party/ unchanged; proto package
  xai.grok.tools.v1 -> kigi.tools.v1
- Config home ~/.kigi (KIGI_SHARE_DIR override), env prefix GROK_* ->
  KIGI_*, `kigi --version` carries the unofficial-community-build notice
- clap identity, help text, startup banner, prompt templates rebranded
  (templates re-encrypted)

Deletions (PRD removal list #5/#6/#7/#9/#10)
- voice input (xai-grok-voice) and all TUI wiring
- telemetry: Mixpanel client, external OTel stream, Sentry, OTLP layers,
  trace/GCS/S3 upload queues (kigi-file-utils halved), workspace upload
  module & dc_log, heap-profile uploader, auth-diagnostics uploader,
  session-analytics halves of feedback; local zero-egress observability
  preserved in new kigi-log crate (unified log, --debug firehose,
  subsystem file logs, opt-in instrumentation)
- announcements (crate, remote-settings fields, TUI surfaces)
- plugin marketplace (crate, sources/browse/CTA/extensions-modal tab);
  direct plugin install/uninstall/update via kigi-agent git_install kept
- relay/gateway/assets endpoints and features (agent relay, headless
  relay transport, gateway bridge, LeaderEnvUrls); leader IPC socket now
  ~/.kigi/leader.sock + KIGI_LEADER_SOCKET, no ws-url derivation
- functional types rehomed instead of deleted: PermissionMode ->
  kigi-config-types, McpInitStrategy -> kigi-mcp, PrCreationSource ->
  session signals, TerminalDiagnostics -> kigi-pager-render, agent_id ->
  shell util

Endpoints
- kigi-env rewritten: single production KigiEndpoints {coding_api_base_url
  https://api.kimi.com/coding/v1 (KIGI_CODE_BASE_URL), oauth_host
  https://auth.kimi.com (KIGI_OAUTH_HOST), update_base_url (GitHub
  Releases API), upgrade_page_url}; GrokBuildEnvironment enum deleted

Toolchain & workspace hygiene
- Rust 1.97.0 pinned; edition 2024; full cargo update; git2 hoisted to
  workspace at 0.21 (Option->Result API migration), quick-xml 0.41
- Root Cargo.toml hand-maintained (PRD §8.1): version 0.1.0 inherited by
  all members, members sorted, unused deps pruned
- cargo-deny advisories gate (deny.toml with documented transitive
  exceptions); CI workflow (check/clippy/fmt/deny/test, macOS+Linux)
- cross-crate test seams re-gated behind `test-support` cargo feature;
  insta snapshot baselines renamed to the kigi_tui prefix
- clippy --workspace --all-targets: zero warnings; fmt clean

Fixes surfaced by the port
- updater probe/installer divergence (bin/kigi vs bin/grok symlink set)
- idle model-metadata refresh dead under KIGI_CODE_BASE_URL override
  (new is_effective_coding_endpoint_url, loopback+override aware)
- macOS symlinked-TMPDIR fixture canonicalization (foreign_sessions,
  fast-worktree); RSS measurement tests serialized via serial_test

Docs & legal (Apache §4)
- NOTICE added (upstream attribution + change statement); THIRD-PARTY
  notices sustained; kigi-tools ported-code notices extended; README,
  CONTRIBUTING, SECURITY, AGENTS.md rewritten

Out of scope for M0 (tracked): Kimi auth/inference (M1), search/fetch,
command parity, config import (M2), Computer Hub excision & final
brand-token sweep (M2), distribution & self-update rewrite (M3).
2026-07-17 05:31:01 -04:00

308 lines
12 KiB
Rust

//! Kill-handle for a session's child-process trees (`Send + Sync`).
//!
//! If a session's worker thread wedges, its `Drop`s never run and child
//! processes leak. A `ProcessScope` lives outside the worker so a supervisor
//! can `kill_all()` one session's children without restarting the host.
//!
//! # Weak-keyed registry (PID-reuse safety)
//!
//! Each child enrolls as a [`ProcessGroup`]. The scope holds only `Weak`
//! references; the spawn-site owner holds the strong `Arc`. On `kill_all`,
//! only groups whose owner is still alive (i.e. un-reaped children) upgrade
//! successfully — reaped children upgrade to `None` and are skipped. This
//! prevents `killpg` from hitting a reused PID.
//!
//! # Residual
//!
//! `kill_all` SIGKILLs but doesn't `wait` (would race the live owner).
//! A wedged owner's killed leader stays as a zombie until the host exits —
//! bounded by wedge events, not sessions.
use std::io;
use std::sync::atomic::{AtomicBool, Ordering};
use std::sync::{Arc, Mutex, OnceLock, PoisonError, Weak};
use crate::{ProcessGroup, new_process_group};
/// A `Send + Sync` kill-handle for one unit's child-process trees. Cheap to
/// clone (shares one inner via `Arc`). See the [module docs](self) for the
/// ownership model and PID-reuse safety argument.
#[derive(Clone)]
pub struct ProcessScope {
inner: Arc<ScopeInner>,
}
struct ScopeInner {
/// One `Weak` per enrolled child tree. The strong `Arc` lives at the spawn
/// site for as long as that child is alive/owned; a dead `Weak` means the
/// owner already reaped+dropped it, so killing it is neither needed nor safe.
groups: Mutex<Vec<Weak<ProcessGroup>>>,
/// Latched once [`kill_all`](ProcessScope::kill_all) has run. Nothing calls
/// `kill_all` twice (the supervisor drops its handle after), so a child
/// enrolled *after* close — a spawn that won the close/spawn race on a
/// wedged actor — would otherwise never be reaped.
/// [`register`](ProcessScope::register) checks this under the `groups` lock
/// and kills such a child on the spot instead of enrolling a leaked `Weak`.
closed: AtomicBool,
}
impl ProcessScope {
/// Create an empty scope. Infallible — group/job handles are created lazily
/// as children are enrolled.
pub fn new() -> Self {
Self {
inner: Arc::new(ScopeInner {
groups: Mutex::new(Vec::new()),
closed: AtomicBool::new(false),
}),
}
}
/// Configure `cmd` so its spawned child becomes the leader of a new process
/// group / job. Call this **before** `cmd.spawn()`, then [`enroll`] the
/// resulting child (or build the group yourself and [`register`] it).
///
/// [`enroll`]: Self::enroll
/// [`register`]: Self::register
pub fn prepare(&self, cmd: &mut tokio::process::Command) {
new_process_group(cmd);
}
/// Register an already-built process group so the scope can reap it if its
/// owner wedges. The scope keeps only a [`Weak`]; **the caller MUST keep the
/// `Arc` alive** for as long as the child is its responsibility. Dropping the
/// last `Arc` (clean reap) makes this registration a silent no-op, which is
/// what keeps [`kill_all`] PID-reuse-safe.
///
/// For spawn sites (e.g. the MCP child handle, the bash terminal) that
/// already build their own [`ProcessGroup`] and own its lifecycle; sites that
/// don't can use [`enroll`] instead.
///
/// [`kill_all`]: Self::kill_all
/// [`enroll`]: Self::enroll
pub fn register(&self, group: &Arc<ProcessGroup>) {
let mut groups = self.lock();
if self.inner.closed.load(Ordering::Relaxed) {
// The scope was already reclaimed (`kill_all` ran) and won't run
// again, so reap this just-spawned child now rather than enroll a
// `Weak` that would leak. Closes the close/spawn race where a
// (possibly wedged) actor's spawn lands after teardown. `killpg` is
// non-blocking, so killing under the lock is fine and serializes
// with a concurrent `kill_all`.
let _ = group.kill();
return;
}
groups.retain(|w| w.strong_count() > 0);
groups.push(Arc::downgrade(group));
}
/// Build a process group for an already-spawned `child` (which must have been
/// configured via [`prepare`]), register a [`Weak`] into this scope, and
/// return the owning `Arc` for the caller to hold. Errors only if the child
/// already exited (nothing to attach) — not a leak.
///
/// [`prepare`]: Self::prepare
#[must_use = "the returned Arc<ProcessGroup> must be kept alive or the scope cannot reap the child"]
pub fn enroll(&self, child: &tokio::process::Child) -> io::Result<Arc<ProcessGroup>> {
let mut group = ProcessGroup::new()?;
group.attach(child)?;
let group = Arc::new(group);
self.register(&group);
Ok(group)
}
/// Convenience for simple sites: [`prepare`] + spawn + [`enroll`]. Returns
/// the child together with the owning `Arc<ProcessGroup>`, which the caller
/// must keep alive for the scope to be able to reap the child.
///
/// [`prepare`]: Self::prepare
/// [`enroll`]: Self::enroll
#[must_use = "the returned Arc<ProcessGroup> must be kept alive or the scope cannot reap the child"]
pub fn spawn(
&self,
mut cmd: tokio::process::Command,
) -> io::Result<(tokio::process::Child, Arc<ProcessGroup>)> {
self.prepare(&mut cmd);
#[allow(clippy::disallowed_methods)]
// ProcessScope::spawn is the enrollment primitive itself.
let child = cmd.spawn()?;
let group = self.enroll(&child)?;
Ok((child, group))
}
/// Idempotently kill every still-owned process tree (`killpg(SIGKILL)` /
/// `TerminateJobObject`). Safe to call multiple times and from any thread.
/// Groups whose owner already reaped+dropped them upgrade to `None` and are
/// skipped — so this never `killpg`s a reused PID.
pub fn kill_all(&self) {
let mut groups = self.lock();
for weak in groups.iter() {
if let Some(group) = weak.upgrade() {
// A failed kill means the group already exited (ESRCH) — benign,
// and there is nothing actionable to log at this primitive layer.
// Callers that care about the wedge-reclaim event log it there.
let _ = group.kill();
}
}
// Every weak has been handled above; clear the set.
groups.clear();
// Latch closed under the lock: a concurrent `register` either already
// pushed (its group was just killed in the loop) or now sees `closed`
// and kills its own child — nothing slips through after teardown.
self.inner.closed.store(true, Ordering::Relaxed);
}
/// Lock the group set, tolerating a poisoned mutex: the critical sections
/// here are panic-free, and a best-effort reaper must still run even if some
/// unrelated thread panicked while holding the lock.
fn lock(&self) -> std::sync::MutexGuard<'_, Vec<Weak<ProcessGroup>>> {
self.inner
.groups
.lock()
.unwrap_or_else(PoisonError::into_inner)
}
/// Number of still-live enrolled groups: weaks whose owning `Arc` has not
/// yet been dropped. This is exactly the set `kill_all` would `killpg`, so a
/// spawn-site owner that has reaped its child (and dropped its `Arc`) is no
/// longer counted — letting callers assert PID-reuse safety.
pub fn live_count(&self) -> usize {
self.lock().iter().filter(|w| w.strong_count() > 0).count()
}
}
impl Default for ProcessScope {
fn default() -> Self {
Self::new()
}
}
/// Process-global [`ProcessScope`] for reaping detached child trees at process
/// exit. Spawn sites (e.g. the local terminal backend) enroll their children
/// here, and the TUI exit paths call [`ProcessScope::kill_all`] on it so a
/// `setsid`-detached background command cannot outlive the process that started
/// it. `kill_all` latches the scope closed, so call it only when the process is
/// genuinely exiting.
pub fn global_process_scope() -> &'static ProcessScope {
static GLOBAL: OnceLock<ProcessScope> = OnceLock::new();
GLOBAL.get_or_init(ProcessScope::new)
}
impl Drop for ScopeInner {
fn drop(&mut self) {
// RAII backstop: if the last scope handle drops without an explicit
// `kill_all`, still reap any group whose owner is alive (a wedged
// unit). Dead weaks (clean teardown) upgrade to None and are skipped,
// so this stays PID-reuse-safe.
let groups = self.groups.lock().unwrap_or_else(PoisonError::into_inner);
for weak in groups.iter() {
if let Some(group) = weak.upgrade() {
let _ = group.kill();
}
}
}
}
#[cfg(all(test, unix))]
mod tests {
use super::*;
use std::time::Duration;
fn sleeper() -> tokio::process::Command {
let mut c = tokio::process::Command::new("sleep");
c.arg("1000");
c
}
/// `wait()` completing == the process actually died (and is reaped). If the
/// kill failed, the `sleep 1000` would run on and `wait()` would time out.
async fn died(child: &mut tokio::process::Child) -> bool {
tokio::time::timeout(Duration::from_secs(3), child.wait())
.await
.is_ok()
}
#[tokio::test]
async fn kill_all_reaps_every_enrolled_child() {
let scope = ProcessScope::new();
// The owner (here, the test) keeps the Arcs alive — as a live spawn site
// would for as long as the child is running.
let (mut c1, _g1) = scope.spawn(sleeper()).unwrap();
let (mut c2, _g2) = scope.spawn(sleeper()).unwrap();
assert_eq!(scope.live_count(), 2);
scope.kill_all();
assert!(died(&mut c1).await, "child 1 must die after kill_all");
assert!(died(&mut c2).await, "child 2 must die after kill_all");
}
#[tokio::test]
async fn kill_all_is_idempotent() {
let scope = ProcessScope::new();
let (mut c, _g) = scope.spawn(sleeper()).unwrap();
scope.kill_all();
scope.kill_all(); // second call must not panic / error
assert!(died(&mut c).await);
}
/// PID-reuse safety: once the owner drops its `Arc` (simulating a clean
/// reap), the scope no longer references the group and `kill_all` is a no-op
/// for it — it must NOT `killpg` a now-reapable/reused group id.
#[tokio::test]
async fn kill_all_skips_group_whose_owner_dropped() {
let scope = ProcessScope::new();
let (mut c, group) = scope.spawn(sleeper()).unwrap();
assert_eq!(scope.live_count(), 1);
// Owner reaps + releases ownership.
drop(group);
assert_eq!(
scope.live_count(),
0,
"dropping the owner's Arc must make the scope's weak dead"
);
scope.kill_all(); // must be a no-op for the now-unowned group
// The child was never killed by the scope; clean it up so the test
// doesn't leak a real `sleep` process.
let _ = c.start_kill();
let _ = c.wait().await;
}
#[tokio::test]
async fn drop_reaps_children_while_owner_alive() {
let scope = ProcessScope::new();
// Owner Arc is held by the test, so the scope's weak is live; dropping
// the scope's last handle must SIGKILL the still-owned group.
let (mut c, _g) = scope.spawn(sleeper()).unwrap();
drop(scope);
assert!(
died(&mut c).await,
"dropping the scope must reap a still-owned enrolled child"
);
}
/// Close/spawn race: a child enrolled *after* `kill_all` must be killed on
/// the spot, not leaked — `kill_all` won't run again to catch it.
#[tokio::test]
async fn register_after_kill_all_reaps_immediately() {
let scope = ProcessScope::new();
scope.kill_all(); // close the scope
let mut cmd = sleeper();
scope.prepare(&mut cmd);
#[allow(clippy::disallowed_methods)] // test: exercises enroll() after close
let mut child = cmd.spawn().unwrap();
let _group = scope.enroll(&child).unwrap(); // register() runs post-close
assert_eq!(
scope.live_count(),
0,
"a post-close register must not enroll the group"
);
assert!(
died(&mut child).await,
"a child registered after kill_all must be killed immediately"
);
}
}