Files
ZacharyZhang-NY 6f31415ed6 §9 acceptance: grep-zero sweep — every internal x.ai/grok identifier renamed
The PRD's first acceptance gate now holds: grep -RinE '\bx\.ai\b|grok'
crates/ --include='*.rs' → 0 matches (exempt: NOTICE and third-party
license archives, README provenance, and the required 'Based on Grok
Build Open Source' attribution, now sourced from version_attribution.txt).

Wire-visible renames (both sides in this repo, changed in lockstep):
- Auth method id 'grok.com' → 'kimi-code' (AuthMethodKind::KimiCode).
- Every x.ai/* and _x.ai/* ACP ext method and meta key → kigi/* /
  _kigi/* (~200 names; grokShell → kigiShell). Session-file replay keeps
  a read-side alias for the legacy '_x.ai/session/update' method so
  existing updates.jsonl histories load; writes emit only the new name
  (both directions test-pinned).
- Agent types grok-build* → kigi* with a documented legacy-prefix alias
  at resolution time so persisted sessions keep resolving.
- ToolNamespace/BuiltinAgentName GrokBuild* → Kigi* (wire snake_case
  kigi/kigi_concise/kigi_hashline; schema regenerated); grok_build
  implementation dirs renamed to kigi*.
- x-grok-* headers → x-kigi-*, __GROK_* sentinels → __KIGI_*, themes
  grokday/groknight → kigiday/kiginight (old persisted values fall back
  to the default theme), web_fetch allowlist xAI hosts → kimi.com +
  moonshot platforms, changelog CDN → this repo, grok-build changelog
  archives deleted.
- BYOK default endpoint removed: [endpoints] api_base_url is now truly
  optional with NO default — consumers fail fast with the flag name when
  unset (no silent x.ai egress). Mock harnesses inject it explicitly.
- System-prompt identity fixed: 'released by xAI' → 'an unofficial
  community CLI for Kimi' (template + regenerated encrypted form).

Also repaired pre-existing grok-era test debt found by the sweep: the
stale trace_classify default-model pin, the grok-pager UA label test,
pty-harness stale-binary reuse and non-hermetic moonshot routing (a PTY
test could previously reach the real api.moonshot.cn), and the outdated
oauth fixture scope key.

Gates: §9 grep 0; fmt clean; workspace check/clippy 0/0 (-D warnings);
FULL cargo test --workspace: 234 suites, 21,961 passed, 0 failed;
deny advisories ok.
2026-07-18 02:48:46 -04:00

97 lines
3.8 KiB
Rust

//! Mock-HTTP integration test for reasoning-only detection on the Responses
//! API wire — the trigger that drives the model doomloop.
//!
//! Spawns a `MockInferenceServer` that serves a `/v1/responses` SSE stream
//! carrying only reasoning (reasoning summary deltas, no output text, no tool
//! call) and asserts the shell sampling client classifies the collected
//! response as `EmptyReason::ReasoningOnly`. This is the exact check the
//! sampler's retry loop runs on every completed turn to decide to resample
//! (and accumulate the out-of-band streaming-capture segments verified by the
//! actor-level capture test).
mod common;
use common::create_test_client;
use kigi_sampling_types::EmptyReason;
use kigi_shell::sampling::{ApiBackend, ConversationItem, ConversationRequest};
use kigi_test_support::sse::responses_api_reasoning_only_events;
use kigi_test_support::{MockInferenceServer, ScriptedResponse};
/// A `/v1/responses` stream that streams only reasoning and finishes with no
/// visible content must be collected into a response the client classifies as
/// `EmptyReason::ReasoningOnly` — the detection that makes the shell resample
/// and spin the doomloop. Exercises the real SSE/HTTP path
/// (`conversation_collect` -> `stream_responses` -> `collect_response`).
#[tokio::test]
async fn responses_api_reasoning_only_is_classified_as_reasoning_only() {
let server = MockInferenceServer::start().await.unwrap();
server.enqueue_response(
"/v1/responses",
ScriptedResponse::sse(responses_api_reasoning_only_events(
"let me think carefully about this",
"kigi-test",
)),
);
let client = create_test_client(&server.url(), ApiBackend::Responses);
let request = ConversationRequest::from_items(vec![ConversationItem::user(
"Solve this without writing anything",
)]);
// The stream completes normally — it is just empty — so collect succeeds.
let response = client
.conversation_collect(request)
.await
.expect("collect must succeed: the stream completes, the response is empty");
assert_eq!(
response.empty_reason(),
Some(EmptyReason::ReasoningOnly),
"a reasoning-only Responses stream must classify as reasoning_only",
);
assert!(response.is_empty());
// The reasoning sibling survived; the assistant is present but empty.
assert!(
response.reasoning_items().next().is_some(),
"the reasoning item must be collected as a sibling",
);
let assistant = response
.assistant()
.expect("an empty assistant is synthesized for the turn");
assert!(
assistant.content.is_empty(),
"reasoning-only means the assistant carried no visible content",
);
}
/// Negative control for the classifier: a normal text `/v1/responses` stream
/// carrying visible assistant content must NOT be classified empty —
/// `empty_reason()` is `None`, distinguishing real content from the
/// reasoning-only case above. (No recovery/resample loop is exercised here.)
#[tokio::test]
async fn normal_text_response_is_not_classified_reasoning_only() {
let server = MockInferenceServer::start().await.unwrap();
server.set_response("The answer is 42.");
let client = create_test_client(&server.url(), ApiBackend::Responses);
let response = client
.conversation_collect(ConversationRequest::from_items(vec![
ConversationItem::user("What is the answer?"),
]))
.await
.unwrap();
assert!(
response.empty_reason().is_none(),
"a normal text turn must not be classified empty",
);
let assistant = response
.assistant()
.expect("assistant present on a text turn");
assert!(
assistant.content.contains("42"),
"the text turn must carry the model's content, got: {:?}",
assistant.content,
);
}