Two T10-flavoured changes in one commit, each independently ship-able
on its own:
1. `web_surface_frame::rgba_hash` switches from std's
`DefaultHasher` (SipHash13, ~1.5 GB/s) to `ahash::AHasher`
(~10 GB/s). At 1080p (8 MB per frame) the dedup key drops from
~5 ms to ~0.8 ms per cache-miss frame, returning roughly 25 % of
the 16 ms scroll budget that was being spent hashing the
newly-arrived RGBA payload. ahash was already in the dependency
graph transitively via hashbrown, so this only adds a direct
`ahash = "0.8"` line and one Cargo.lock entry.
2. `docs/t10-iosurface-plan.md` records the full architectural
roadmap for the actual zero-copy path that supersedes the
software-pipe pipeline: `OffscreenRenderingContext` against a
hardware surfman adapter, IOSurface-backed surface on macOS,
mach-port handoff to the GPUI process, MTLTexture import as an
external sampler. The document explains why each currently
shipped commit (`840255f`, `a80d039`, `e02c0fd`, `7f3b8b4`, plus
this hash swap) is a stepping stone that eventually deletes
itself once the IOSurface path lands, and names the upstream
API gap in `servo-paint-api` that blocks step 2.
cargo test --bin ely_app: 120 passed, 0 failed, 2 ignored.
Hash collision probability remains ~1 in 2^64; AHash uses the same
keyspace as the previous SipHash13.
WebSurfaceFrame::from_parts was unconditionally calling
Arc::new(RenderImage::new([image::Frame::new(image_buffer)])) on
every live frame, even when the underlying bytes were
byte-for-byte identical to the previous frame. At 60 fps on a 1080p
canvas that was ~960 MB/s of host-side cloning plus a fresh GPUI
texture allocation on every tick — the bottleneck Linus + Karpathy
+ Jony flagged as the next material step after dropping the
file-system pixel pipe (a80d039).
A thread_local single-slot cache in web_surface_frame.rs now keys
on a 64-bit DefaultHasher of the raw RGBA bytes. On a cache hit
the existing Arc<RenderImage> is reused; on a miss the buffer is
built once, stored, and returned. Steady-state idle pages stop
churning the GPUI texture pool entirely. Hash collisions are
1 in 2^64 — if that ever becomes a real worry, the cache key can
be widened to length + a sample of bytes before paying the full
memcmp; not worth doing today.
The T10 red guard
(identical_live_frames_share_render_image_arc) drops its
#[ignore] attribute outright per the contract it documented.
The T7 red guard (user_click_in_rendered_web_canvas_reaches_input_pipeline)
remains ignored — it's a separate diagnosis tracked under T13.
cargo test --bin ely_app: 113 passed, 0 failed, 1 ignored
(was 112+2, T10 guard went green).
cargo test --bin ely_app -- --ignored: 1 failed (only T7 click
pipeline remains red).
Follow-ups (left for the endgame T10 IOSurface path):
* a per-tab cache would prevent multi-tab switching from
thrashing the single slot; defer until a real multi-tab
scroll benchmark shows it matters.
* the endgame is OffscreenRenderingContext + IOSurface so the
GPU texture itself is the source of truth and the host-side
Vec<u8> + ImageBuffer + RenderImage allocation chain
disappears entirely.