The PRD's first acceptance gate now holds: grep -RinE '\bx\.ai\b|grok' crates/ --include='*.rs' → 0 matches (exempt: NOTICE and third-party license archives, README provenance, and the required 'Based on Grok Build Open Source' attribution, now sourced from version_attribution.txt). Wire-visible renames (both sides in this repo, changed in lockstep): - Auth method id 'grok.com' → 'kimi-code' (AuthMethodKind::KimiCode). - Every x.ai/* and _x.ai/* ACP ext method and meta key → kigi/* / _kigi/* (~200 names; grokShell → kigiShell). Session-file replay keeps a read-side alias for the legacy '_x.ai/session/update' method so existing updates.jsonl histories load; writes emit only the new name (both directions test-pinned). - Agent types grok-build* → kigi* with a documented legacy-prefix alias at resolution time so persisted sessions keep resolving. - ToolNamespace/BuiltinAgentName GrokBuild* → Kigi* (wire snake_case kigi/kigi_concise/kigi_hashline; schema regenerated); grok_build implementation dirs renamed to kigi*. - x-grok-* headers → x-kigi-*, __GROK_* sentinels → __KIGI_*, themes grokday/groknight → kigiday/kiginight (old persisted values fall back to the default theme), web_fetch allowlist xAI hosts → kimi.com + moonshot platforms, changelog CDN → this repo, grok-build changelog archives deleted. - BYOK default endpoint removed: [endpoints] api_base_url is now truly optional with NO default — consumers fail fast with the flag name when unset (no silent x.ai egress). Mock harnesses inject it explicitly. - System-prompt identity fixed: 'released by xAI' → 'an unofficial community CLI for Kimi' (template + regenerated encrypted form). Also repaired pre-existing grok-era test debt found by the sweep: the stale trace_classify default-model pin, the grok-pager UA label test, pty-harness stale-binary reuse and non-hermetic moonshot routing (a PTY test could previously reach the real api.moonshot.cn), and the outdated oauth fixture scope key. Gates: §9 grep 0; fmt clean; workspace check/clippy 0/0 (-D warnings); FULL cargo test --workspace: 234 suites, 21,961 passed, 0 failed; deny advisories ok.
97 lines
3.8 KiB
Rust
97 lines
3.8 KiB
Rust
//! Mock-HTTP integration test for reasoning-only detection on the Responses
|
|
//! API wire — the trigger that drives the model doomloop.
|
|
//!
|
|
//! Spawns a `MockInferenceServer` that serves a `/v1/responses` SSE stream
|
|
//! carrying only reasoning (reasoning summary deltas, no output text, no tool
|
|
//! call) and asserts the shell sampling client classifies the collected
|
|
//! response as `EmptyReason::ReasoningOnly`. This is the exact check the
|
|
//! sampler's retry loop runs on every completed turn to decide to resample
|
|
//! (and accumulate the out-of-band streaming-capture segments verified by the
|
|
//! actor-level capture test).
|
|
|
|
mod common;
|
|
|
|
use common::create_test_client;
|
|
use kigi_sampling_types::EmptyReason;
|
|
use kigi_shell::sampling::{ApiBackend, ConversationItem, ConversationRequest};
|
|
use kigi_test_support::sse::responses_api_reasoning_only_events;
|
|
use kigi_test_support::{MockInferenceServer, ScriptedResponse};
|
|
|
|
/// A `/v1/responses` stream that streams only reasoning and finishes with no
|
|
/// visible content must be collected into a response the client classifies as
|
|
/// `EmptyReason::ReasoningOnly` — the detection that makes the shell resample
|
|
/// and spin the doomloop. Exercises the real SSE/HTTP path
|
|
/// (`conversation_collect` -> `stream_responses` -> `collect_response`).
|
|
#[tokio::test]
|
|
async fn responses_api_reasoning_only_is_classified_as_reasoning_only() {
|
|
let server = MockInferenceServer::start().await.unwrap();
|
|
server.enqueue_response(
|
|
"/v1/responses",
|
|
ScriptedResponse::sse(responses_api_reasoning_only_events(
|
|
"let me think carefully about this",
|
|
"kigi-test",
|
|
)),
|
|
);
|
|
let client = create_test_client(&server.url(), ApiBackend::Responses);
|
|
|
|
let request = ConversationRequest::from_items(vec![ConversationItem::user(
|
|
"Solve this without writing anything",
|
|
)]);
|
|
|
|
// The stream completes normally — it is just empty — so collect succeeds.
|
|
let response = client
|
|
.conversation_collect(request)
|
|
.await
|
|
.expect("collect must succeed: the stream completes, the response is empty");
|
|
|
|
assert_eq!(
|
|
response.empty_reason(),
|
|
Some(EmptyReason::ReasoningOnly),
|
|
"a reasoning-only Responses stream must classify as reasoning_only",
|
|
);
|
|
assert!(response.is_empty());
|
|
// The reasoning sibling survived; the assistant is present but empty.
|
|
assert!(
|
|
response.reasoning_items().next().is_some(),
|
|
"the reasoning item must be collected as a sibling",
|
|
);
|
|
let assistant = response
|
|
.assistant()
|
|
.expect("an empty assistant is synthesized for the turn");
|
|
assert!(
|
|
assistant.content.is_empty(),
|
|
"reasoning-only means the assistant carried no visible content",
|
|
);
|
|
}
|
|
|
|
/// Negative control for the classifier: a normal text `/v1/responses` stream
|
|
/// carrying visible assistant content must NOT be classified empty —
|
|
/// `empty_reason()` is `None`, distinguishing real content from the
|
|
/// reasoning-only case above. (No recovery/resample loop is exercised here.)
|
|
#[tokio::test]
|
|
async fn normal_text_response_is_not_classified_reasoning_only() {
|
|
let server = MockInferenceServer::start().await.unwrap();
|
|
server.set_response("The answer is 42.");
|
|
let client = create_test_client(&server.url(), ApiBackend::Responses);
|
|
|
|
let response = client
|
|
.conversation_collect(ConversationRequest::from_items(vec![
|
|
ConversationItem::user("What is the answer?"),
|
|
]))
|
|
.await
|
|
.unwrap();
|
|
|
|
assert!(
|
|
response.empty_reason().is_none(),
|
|
"a normal text turn must not be classified empty",
|
|
);
|
|
let assistant = response
|
|
.assistant()
|
|
.expect("assistant present on a text turn");
|
|
assert!(
|
|
assistant.content.contains("42"),
|
|
"the text turn must carry the model's content, got: {:?}",
|
|
assistant.content,
|
|
);
|
|
}
|