F3: Kimi inference pipeline + full grok cloud-surface excision

Sampler / inference (PRD F3):
- kimi_compat.rs: single adaptation point for the Kimi chat/completions
  dialect (thinking-field mapping, model_id stripping, empty-content
  tool-call message fix, stream_options.include_usage), with kimi-cli
  source citations
- Rate-limit handling reworked for Kimi/Moonshot semantics; UA kigi/{version}
- /models replaces the xAI models-v2 endpoint everywhere; idle model
  refresh carries X-Msh-* device headers only (X-XAI-Token-Auth and
  x-grok-client-mode/CLIENT_MODE_HEADER machinery deleted)

Cloud-surface excision (PRD §5, zero-egress):
- remote/ conversations lane, cli-chat-proxy-types crate, prod/ dir,
  share command, credit bar: deleted (single local session lane;
  paginate() replaces merge_and_paginate)
- Subscription/tier gate stack deleted end-to-end: AppView
  gate/tier/team/ZDR fields, app/subscription.rs watch loop,
  dispatch/billing.rs paywall + SuperGrok upsell, free-usage-exhausted
  chain, tier-restricted commands, GateInfo, RemoteSettings gate fields,
  SettingsUpdateNotification gate fields
- /privacy + coding-data-sharing setting deleted (backed by a dead xAI
  RPC; Kigi is zero-egress — nothing to share or retain remotely)

Auth UX correctness (user-reported):
- Device-flow fixtures now mirror the live Kimi payload shape
  (https://www.kimi.com/code/authorize_device?user_code=..., verified
  against auth.kimi.com); the fabricated auth.kimi.com/device?code=...
  URLs are gone
- open_browser_detached is a no-op under cfg(test): unit tests drove
  wiremock fixture URLs into the real browser (root cause of the
  "garbage mock link" ABCD-1234 tabs)
- Welcome/pager-minimal rebrand: Grok Build -> Kigi, grok.com ->
  kimi.com, "Sign in to Grok" -> "Sign in to Kimi"
This commit is contained in:
2026-07-17 16:05:51 -04:00
parent fe1f885bb3
commit ea0ce9d15f
231 changed files with 4730 additions and 26358 deletions
@@ -373,16 +373,13 @@ async fn apply_retry_decision(
RetryDecision::Fatal(fatal_err) => {
// Emit only on true budget exhaustion (hit the retry / rate-limit
// cap), mirroring `classify_error`'s Fatal conditions — NOT on a
// server `x-should-retry: false` or a non-retryable error, which
// are also Fatal but are not "exhausted".
// non-retryable error, which is also Fatal but not "exhausted".
let next_attempt = *retry_count + 1;
let server_said_stop = matches!(err.should_retry_header(), Some(false));
let budget_exhausted = !server_said_stop
&& if err.is_rate_limited() {
next_attempt >= max_retries.min(rate_limit_threshold)
} else {
err.is_retryable() && next_attempt >= max_retries
};
let budget_exhausted = if err.is_rate_limited() {
next_attempt >= max_retries.min(rate_limit_threshold)
} else {
err.is_retryable() && next_attempt >= max_retries
};
if budget_exhausted {
let exhausted_span = tracing::info_span!(
"http.retries_exhausted",
@@ -609,7 +606,6 @@ fn synthesize_from_info(info: &SamplingErrorInfo) -> SamplingError {
message: info.message.clone(),
model_metadata: info.model_metadata.clone(),
retry_after_secs: info.retry_after_secs,
should_retry: None,
}
}
SamplingErrorKind::EmptyResponse => {