Files
Kigi-CLI/crates/codegen/kigi-tools/tests/cgroup_memory_test.rs
T
ZacharyZhang-NY d6c20fc13f M0: compilable skeleton — Kigi 0.1.0 fork surgery
Hard fork of xai-org/grok-build (Apache-2.0) re-targeted as Kigi, an
unofficial Kimi Code CLI community build.

Rename & identity
- 72 xai-*/xai-grok-* crates -> kigi-* (explicit: xai-grok-pager-bin ->
  kigi-bin [binary `kigi`], xai-grok-pager -> kigi-tui; rest mechanical);
  ptyctl, ptyctl-cli, third_party/ unchanged; proto package
  xai.grok.tools.v1 -> kigi.tools.v1
- Config home ~/.kigi (KIGI_SHARE_DIR override), env prefix GROK_* ->
  KIGI_*, `kigi --version` carries the unofficial-community-build notice
- clap identity, help text, startup banner, prompt templates rebranded
  (templates re-encrypted)

Deletions (PRD removal list #5/#6/#7/#9/#10)
- voice input (xai-grok-voice) and all TUI wiring
- telemetry: Mixpanel client, external OTel stream, Sentry, OTLP layers,
  trace/GCS/S3 upload queues (kigi-file-utils halved), workspace upload
  module & dc_log, heap-profile uploader, auth-diagnostics uploader,
  session-analytics halves of feedback; local zero-egress observability
  preserved in new kigi-log crate (unified log, --debug firehose,
  subsystem file logs, opt-in instrumentation)
- announcements (crate, remote-settings fields, TUI surfaces)
- plugin marketplace (crate, sources/browse/CTA/extensions-modal tab);
  direct plugin install/uninstall/update via kigi-agent git_install kept
- relay/gateway/assets endpoints and features (agent relay, headless
  relay transport, gateway bridge, LeaderEnvUrls); leader IPC socket now
  ~/.kigi/leader.sock + KIGI_LEADER_SOCKET, no ws-url derivation
- functional types rehomed instead of deleted: PermissionMode ->
  kigi-config-types, McpInitStrategy -> kigi-mcp, PrCreationSource ->
  session signals, TerminalDiagnostics -> kigi-pager-render, agent_id ->
  shell util

Endpoints
- kigi-env rewritten: single production KigiEndpoints {coding_api_base_url
  https://api.kimi.com/coding/v1 (KIGI_CODE_BASE_URL), oauth_host
  https://auth.kimi.com (KIGI_OAUTH_HOST), update_base_url (GitHub
  Releases API), upgrade_page_url}; GrokBuildEnvironment enum deleted

Toolchain & workspace hygiene
- Rust 1.97.0 pinned; edition 2024; full cargo update; git2 hoisted to
  workspace at 0.21 (Option->Result API migration), quick-xml 0.41
- Root Cargo.toml hand-maintained (PRD §8.1): version 0.1.0 inherited by
  all members, members sorted, unused deps pruned
- cargo-deny advisories gate (deny.toml with documented transitive
  exceptions); CI workflow (check/clippy/fmt/deny/test, macOS+Linux)
- cross-crate test seams re-gated behind `test-support` cargo feature;
  insta snapshot baselines renamed to the kigi_tui prefix
- clippy --workspace --all-targets: zero warnings; fmt clean

Fixes surfaced by the port
- updater probe/installer divergence (bin/kigi vs bin/grok symlink set)
- idle model-metadata refresh dead under KIGI_CODE_BASE_URL override
  (new is_effective_coding_endpoint_url, loopback+override aware)
- macOS symlinked-TMPDIR fixture canonicalization (foreign_sessions,
  fast-worktree); RSS measurement tests serialized via serial_test

Docs & legal (Apache §4)
- NOTICE added (upstream attribution + change statement); THIRD-PARTY
  notices sustained; kigi-tools ported-code notices extended; README,
  CONTRIBUTING, SECURITY, AGENTS.md rewritten

Out of scope for M0 (tracked): Kimi auth/inference (M1), search/fetch,
command parity, config import (M2), Computer Hub excision & final
brand-token sweep (M2), distribution & self-update rewrite (M3).
2026-07-17 05:31:01 -04:00

426 lines
14 KiB
Rust
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
//! Integration test for cgroup memory-high OOM handling.
//!
//! **Must be run on Linux with cgroupv2** and sufficient permissions to create
//! child cgroups (typically root, or a user-session cgroup with delegation).
//!
//! Run with:
//! ```bash
//! # On a Linux machine (as root or with cgroup delegation):
//! cargo test -p kigi-tools --test cgroup_memory_test -- --ignored --nocapture
//!
//! # If you need root:
//! sudo -E cargo test -p kigi-tools --test cgroup_memory_test -- --ignored --nocapture
//! ```
//!
//! The cgroup-dependent tests (15) are `#[ignore]`d by default so they don't
//! run in CI where cgroup delegation is typically unavailable. Test 6 (no-config)
//! always runs.
//!
//! The tests exercise:
//! 1. A command that stays under the memory limit → exits normally (exit 0)
//! 2. A command that exceeds memory.high → killed with exit 137, signal "oom"
//! 3. The session (backend) survives an OOM and can run another command after
//! 4. Background tasks are also killed on OOM
//! 5. A gradual allocator that slowly ramps up past the limit
use std::collections::HashMap;
use std::path::PathBuf;
use std::time::Duration;
use kigi_tools::computer::local::LocalTerminalBackend;
use kigi_tools::computer::local::cgroup::{CgroupMemoryConfig, PROCESS_OOM_EXIT_CODE};
use kigi_tools::computer::types::{TerminalBackend, TerminalRunRequest, TerminalRunResult};
use kigi_tools::notification::types::ToolNotificationHandle;
// ── Helpers ──────────────────────────────────────────────────────────────
/// Small memory limit for testing: 32 MiB high, 32 MiB headroom (64 MiB hard max).
fn test_memory_config() -> CgroupMemoryConfig {
CgroupMemoryConfig {
memory_high_bytes: 32 * 1024 * 1024, // 32 MiB
headroom_bytes: 32 * 1024 * 1024, // 32 MiB headroom → 64 MiB hard max
}
}
fn make_request(command: &str, timeout_secs: u64) -> TerminalRunRequest {
let output_file = std::env::temp_dir().join(format!(
"cgroup-test-{}-{}.out",
std::process::id(),
std::time::SystemTime::now()
.duration_since(std::time::UNIX_EPOCH)
.unwrap()
.as_nanos()
));
TerminalRunRequest {
command: command.to_string(),
working_directory: PathBuf::from("/tmp"),
env: HashMap::new(),
timeout: Duration::from_secs(timeout_secs),
output_byte_limit: 1024 * 1024,
output_file,
notification_handle: ToolNotificationHandle::noop(),
tool_call_id: format!("cgroup-test-{}", uuid::Uuid::now_v7()),
display_command: None,
auto_background_on_timeout: false,
foreground_block_budget: None,
kind: Default::default(),
owner_session_id: None,
}
}
fn is_linux_with_cgroupv2() -> bool {
#[cfg(target_os = "linux")]
{
// Check that cgroupv2 is mounted
std::path::Path::new("/sys/fs/cgroup/cgroup.controllers").exists()
}
#[cfg(not(target_os = "linux"))]
{
false
}
}
fn can_create_cgroups() -> bool {
if !is_linux_with_cgroupv2() {
return false;
}
// Try reading our own cgroup path — if this works, we can probably create children
#[cfg(target_os = "linux")]
{
if let Ok(contents) = std::fs::read_to_string("/proc/self/cgroup") {
for line in contents.lines() {
if let Some(path) = line.strip_prefix("0::") {
let cgroup_dir = std::path::PathBuf::from(format!("/sys/fs/cgroup{}", path));
// Check if we can write to this cgroup's subtree_control
let subtree = cgroup_dir.join("cgroup.subtree_control");
return subtree.exists();
}
}
}
false
}
#[cfg(not(target_os = "linux"))]
{
false
}
}
fn skip_unless_cgroup() {
if !can_create_cgroups() {
eprintln!(
"\n╔══════════════════════════════════════════════════════════════╗\n\
║ SKIPPED: cgroupv2 not available or insufficient perms. ║\n\
║ Run on Linux as root or with cgroup delegation. ║\n\
╚══════════════════════════════════════════════════════════════╝\n"
);
}
}
fn print_result(label: &str, result: &TerminalRunResult) {
let output_preview = if result.combined_output.len() > 200 {
format!("{}…", &result.combined_output[..200])
} else {
result.combined_output.clone()
};
eprintln!(
"\n── {label} ──\n exit_code: {:?}\n signal: {:?}\n timed_out: {}\n truncated: {}\n output: {:?}\n",
result.exit_code,
result.signal,
result.timed_out,
result.truncated,
output_preview.trim(),
);
}
// ── Tests ────────────────────────────────────────────────────────────────
/// Test 1: A command that stays well under the limit exits normally.
#[tokio::test]
#[ignore = "requires Linux cgroupv2 with delegation — run with: cargo test --test cgroup_memory_test -- --ignored --nocapture"]
async fn test_under_limit_exits_normally() {
skip_unless_cgroup();
if !can_create_cgroups() {
return;
}
eprintln!("\n=== Test: under_limit_exits_normally ===");
let backend = LocalTerminalBackend::with_memory_limit(test_memory_config());
// Give actor time to initialize cgroup
tokio::time::sleep(Duration::from_millis(200)).await;
let result = backend
.run(make_request(
"echo 'hello from cgroup'; cat /proc/self/cgroup",
10,
))
.await
.expect("command should succeed");
print_result("Under limit", &result);
assert_eq!(result.exit_code, Some(0), "Expected exit code 0");
assert!(
result.combined_output.contains("hello from cgroup"),
"Output should contain our echo"
);
assert_ne!(
result.signal.as_deref(),
Some("oom"),
"Should NOT be OOM-killed"
);
eprintln!("✅ PASSED: under_limit_exits_normally");
}
/// Test 2: A command that allocates way more than the limit is killed with 137/oom.
#[tokio::test]
#[ignore = "requires Linux cgroupv2 with delegation"]
async fn test_over_limit_gets_oom_killed() {
skip_unless_cgroup();
if !can_create_cgroups() {
return;
}
eprintln!("\n=== Test: over_limit_gets_oom_killed ===");
let backend = LocalTerminalBackend::with_memory_limit(test_memory_config());
tokio::time::sleep(Duration::from_millis(200)).await;
// Allocate 128 MiB in Python — well above the 32 MiB high / 64 MiB max limits.
let alloc_cmd = r#"python3 -c "
import sys
print('Allocating 128 MiB...', flush=True)
data = bytearray(128 * 1024 * 1024)
print('Allocation succeeded (should not reach here)', flush=True)
""#;
let result = backend
.run(make_request(alloc_cmd, 30))
.await
.expect("command should return a result (even if killed)");
print_result("Over limit", &result);
// The process should be killed — either by our monitor (exit 137 + signal "oom")
// or by the kernel hard OOM killer (exit 137 / signal 9).
let killed_by_memory = result.exit_code == Some(PROCESS_OOM_EXIT_CODE)
|| result.signal.as_deref() == Some("oom")
|| result
.signal
.as_ref()
.is_some_and(|s| s.contains("signal 9"));
assert!(
killed_by_memory,
"Expected OOM kill (exit 137 or signal 9/oom), got exit_code={:?} signal={:?}",
result.exit_code, result.signal
);
// Output before the kill should be preserved
assert!(
result.combined_output.contains("Allocating 128 MiB"),
"Output before OOM should be captured"
);
eprintln!("✅ PASSED: over_limit_gets_oom_killed");
}
/// Test 3: After an OOM, the backend still works for subsequent commands.
#[tokio::test]
#[ignore = "requires Linux cgroupv2 with delegation"]
async fn test_session_survives_oom() {
skip_unless_cgroup();
if !can_create_cgroups() {
return;
}
eprintln!("\n=== Test: session_survives_oom ===");
let backend = LocalTerminalBackend::with_memory_limit(test_memory_config());
tokio::time::sleep(Duration::from_millis(200)).await;
// First: trigger an OOM
let oom_cmd = r#"python3 -c "data = bytearray(128 * 1024 * 1024)""#;
let oom_result = backend
.run(make_request(oom_cmd, 30))
.await
.expect("should return result even on OOM");
print_result("OOM command", &oom_result);
// Small delay so cgroup memory is reclaimed
tokio::time::sleep(Duration::from_millis(500)).await;
// Second: run a lightweight command — should succeed
let ok_result = backend
.run(make_request("echo 'alive after OOM'", 10))
.await
.expect("post-OOM command should succeed");
print_result("After OOM", &ok_result);
assert_eq!(
ok_result.exit_code,
Some(0),
"Post-OOM command should exit 0"
);
assert!(
ok_result.combined_output.contains("alive after OOM"),
"Post-OOM output should contain our echo"
);
eprintln!("✅ PASSED: session_survives_oom");
}
/// Test 4: Background tasks are also subject to the memory limit.
#[tokio::test]
#[ignore = "requires Linux cgroupv2 with delegation"]
async fn test_background_task_oom() {
skip_unless_cgroup();
if !can_create_cgroups() {
return;
}
eprintln!("\n=== Test: background_task_oom ===");
let backend = LocalTerminalBackend::with_memory_limit(test_memory_config());
tokio::time::sleep(Duration::from_millis(200)).await;
// Start a background command that will OOM
let alloc_cmd = r#"python3 -c "
import time
print('BG: allocating...', flush=True)
time.sleep(0.5)
data = bytearray(128 * 1024 * 1024)
print('BG: done (should not reach here)', flush=True)
time.sleep(60)
""#;
let handle = backend
.run_background(make_request(alloc_cmd, 60))
.await
.expect("background spawn should succeed");
eprintln!(" Background task_id: {}", handle.task_id);
// Wait for completion (it should be killed before the 60s timeout)
let snapshot = backend
.wait_for_completion(&handle.task_id, Some(Duration::from_secs(30)))
.await;
if let Some(snap) = &snapshot {
eprintln!(
" BG result: completed={} exit_code={:?} signal={:?} output={:?}",
snap.completed,
snap.exit_code,
snap.signal,
&snap.output[..snap.output.len().min(200)]
);
assert!(snap.completed, "Background task should have completed");
let killed_by_memory = snap.exit_code == Some(PROCESS_OOM_EXIT_CODE)
|| snap.signal.as_deref() == Some("oom")
|| snap.signal.as_ref().is_some_and(|s| s.contains("signal 9"));
assert!(
killed_by_memory,
"Background task should be OOM-killed, got exit_code={:?} signal={:?}",
snap.exit_code, snap.signal
);
} else {
panic!("Expected a snapshot for the background task");
}
eprintln!("✅ PASSED: background_task_oom");
}
/// Test 5: Gradual allocation that slowly ramps past the limit.
/// This tests that the inotify monitor catches the memory.high event
/// rather than relying on the kernel's hard memory.max kill.
#[tokio::test]
#[ignore = "requires Linux cgroupv2 with delegation"]
async fn test_gradual_allocation_oom() {
skip_unless_cgroup();
if !can_create_cgroups() {
return;
}
eprintln!("\n=== Test: gradual_allocation_oom ===");
let backend = LocalTerminalBackend::with_memory_limit(test_memory_config());
tokio::time::sleep(Duration::from_millis(200)).await;
// Allocate in 1 MiB chunks with a small delay — slowly ramps past 32 MiB.
let gradual_cmd = r#"python3 -c "
import time, sys
chunks = []
for i in range(128):
chunks.append(bytearray(1024 * 1024)) # 1 MiB per chunk
print(f'Allocated {i+1} MiB', flush=True)
time.sleep(0.05)
print('Finished all allocations (should not reach here)', flush=True)
""#;
let result = backend
.run(make_request(gradual_cmd, 30))
.await
.expect("should return result");
print_result("Gradual allocation", &result);
let killed_by_memory = result.exit_code == Some(PROCESS_OOM_EXIT_CODE)
|| result.signal.as_deref() == Some("oom")
|| result
.signal
.as_ref()
.is_some_and(|s| s.contains("signal 9"));
assert!(
killed_by_memory,
"Expected OOM kill for gradual allocator, got exit_code={:?} signal={:?}",
result.exit_code, result.signal
);
// Should have some output showing allocations before the kill
assert!(
result.combined_output.contains("Allocated"),
"Should see some allocation progress before kill"
);
// Should NOT have finished all 128 MiB
assert!(
!result.combined_output.contains("Finished all allocations"),
"Should have been killed before finishing"
);
eprintln!("✅ PASSED: gradual_allocation_oom");
}
/// Test 6: No memory config → no cgroup enforcement, large alloc succeeds.
/// This verifies the no-op path works correctly.
#[tokio::test]
async fn test_no_config_no_enforcement() {
eprintln!("\n=== Test: no_config_no_enforcement ===");
// Use the plain `new()` constructor — no memory limits
let backend = LocalTerminalBackend::new();
tokio::time::sleep(Duration::from_millis(200)).await;
// Allocate 64 MiB — would be killed with a 32 MiB limit, but should succeed here
let alloc_cmd = r#"python3 -c "
data = bytearray(64 * 1024 * 1024)
print('Allocated 64 MiB without limits')
""#;
let result = backend
.run(make_request(alloc_cmd, 10))
.await
.expect("command should succeed without limits");
print_result("No enforcement", &result);
assert_eq!(result.exit_code, Some(0), "Should exit 0 without limits");
assert!(
result.combined_output.contains("Allocated 64 MiB"),
"Allocation should succeed without limits"
);
eprintln!("✅ PASSED: no_config_no_enforcement");
}