F3: Kimi inference pipeline + full grok cloud-surface excision
Sampler / inference (PRD F3):
- kimi_compat.rs: single adaptation point for the Kimi chat/completions
dialect (thinking-field mapping, model_id stripping, empty-content
tool-call message fix, stream_options.include_usage), with kimi-cli
source citations
- Rate-limit handling reworked for Kimi/Moonshot semantics; UA kigi/{version}
- /models replaces the xAI models-v2 endpoint everywhere; idle model
refresh carries X-Msh-* device headers only (X-XAI-Token-Auth and
x-grok-client-mode/CLIENT_MODE_HEADER machinery deleted)
Cloud-surface excision (PRD §5, zero-egress):
- remote/ conversations lane, cli-chat-proxy-types crate, prod/ dir,
share command, credit bar: deleted (single local session lane;
paginate() replaces merge_and_paginate)
- Subscription/tier gate stack deleted end-to-end: AppView
gate/tier/team/ZDR fields, app/subscription.rs watch loop,
dispatch/billing.rs paywall + SuperGrok upsell, free-usage-exhausted
chain, tier-restricted commands, GateInfo, RemoteSettings gate fields,
SettingsUpdateNotification gate fields
- /privacy + coding-data-sharing setting deleted (backed by a dead xAI
RPC; Kigi is zero-egress — nothing to share or retain remotely)
Auth UX correctness (user-reported):
- Device-flow fixtures now mirror the live Kimi payload shape
(https://www.kimi.com/code/authorize_device?user_code=..., verified
against auth.kimi.com); the fabricated auth.kimi.com/device?code=...
URLs are gone
- open_browser_detached is a no-op under cfg(test): unit tests drove
wiremock fixture URLs into the real browser (root cause of the
"garbage mock link" ABCD-1234 tabs)
- Welcome/pager-minimal rebrand: Grok Build -> Kigi, grok.com ->
kimi.com, "Sign in to Grok" -> "Sign in to Kimi"
This commit is contained in:
@@ -373,16 +373,13 @@ async fn apply_retry_decision(
|
||||
RetryDecision::Fatal(fatal_err) => {
|
||||
// Emit only on true budget exhaustion (hit the retry / rate-limit
|
||||
// cap), mirroring `classify_error`'s Fatal conditions — NOT on a
|
||||
// server `x-should-retry: false` or a non-retryable error, which
|
||||
// are also Fatal but are not "exhausted".
|
||||
// non-retryable error, which is also Fatal but not "exhausted".
|
||||
let next_attempt = *retry_count + 1;
|
||||
let server_said_stop = matches!(err.should_retry_header(), Some(false));
|
||||
let budget_exhausted = !server_said_stop
|
||||
&& if err.is_rate_limited() {
|
||||
next_attempt >= max_retries.min(rate_limit_threshold)
|
||||
} else {
|
||||
err.is_retryable() && next_attempt >= max_retries
|
||||
};
|
||||
let budget_exhausted = if err.is_rate_limited() {
|
||||
next_attempt >= max_retries.min(rate_limit_threshold)
|
||||
} else {
|
||||
err.is_retryable() && next_attempt >= max_retries
|
||||
};
|
||||
if budget_exhausted {
|
||||
let exhausted_span = tracing::info_span!(
|
||||
"http.retries_exhausted",
|
||||
@@ -609,7 +606,6 @@ fn synthesize_from_info(info: &SamplingErrorInfo) -> SamplingError {
|
||||
message: info.message.clone(),
|
||||
model_metadata: info.model_metadata.clone(),
|
||||
retry_after_secs: info.retry_after_secs,
|
||||
should_retry: None,
|
||||
}
|
||||
}
|
||||
SamplingErrorKind::EmptyResponse => {
|
||||
|
||||
Reference in New Issue
Block a user