Files
llama.cpp/tools/ui/tests/client/README-perf.md
T
parabelboi 4dd127584b ui: add read_media tool (#25877)
* server: add read_image tool (#25875)

Adds a server-tool that allows vision models to analyze server-side images.
This tool is reading a single file for now:
The image data is base64 encoded and passed to the UI, which
decodes it, fills the <img> tag and removes the data URI before
passing the tool result back to the model.

* cleanup read_image tool: move magic strings to constants

* Add dedicated constants file: tools/ui/src/lib/constants/read-image.ts
  with PREFIX_IMAGE, PREFIX_SIZE, PREFIX_MIME constants
* Use ATTACHMENT_SAVED_REGEX from agentic.ts in ChatMessageToolCallBlockReadImage.svelte
* Use NEWLINE constant from code.ts instead of hardcoded '\n'
* Use PREFIX_SIZE in regex pattern for size parsing
* Add SERVER_TOOL_READ_IMAGE_PREFIX_* constants in C++ server-tools.cpp
  to match the TypeScript PREFIX_* constants for consistency

* server: rename read_image tool to read_media for images and audio

* Rename server_tool_read_image to server_tool_read_media in C++
* Rename enum BuiltInTool.READ_IMAGE to READ_MEDIA
* Rename UI constants, parser, and Svelte component files
* Update display label from 'Read image' to 'Read media'

* ui: consolidate audio data URI handling into shared utility

* Extract getAudioInputFormat to a shared utility (was duplicated inline)
* Store raw base64 in base64Data on the message object
* Use base64Data to construct data URIs for audio rendering
* Update agentic store to build INPUT_AUDIO parts from base64Data

* server: read_media: restrict audio to wav/mp3 and minor fixes

* Server get_mime_from_extension now only advertises audio/wav and
  audio/mpeg (the only formats the model's input_audio API accepts)
* Case-insensitive extension matching (fixes .MP3, .Wav, etc.)
* Unknown extensions return an error instead of a multi-MB data URI
  that inflates model context with garbage
* Updated tool description to document supported formats
* Frontend AUDIO_MIME_TO_EXTENSION trimmed to match server
* fix a missing import in tools/ui/src/lib/stores/agentic.svelte.ts

* server: read_media: add to --tools help text and README tool list

* ui: fix indentation in ChatMessageToolCallBlockDefault.svelte

* server: read_media tool: fix a cast to use the correct type

* server: read_media: multiple fixes

* server-tools.cpp import cctype, remove UTF-8 char, check mime before reading file
* ui: add MimeTypePrefix.AUDIO and use it in agentic.svelte.ts

* server: make read_media inherit from read_file and add uses_cwd

* ui: fix formating issues

* rm from server

* move it to frontend-only tool

* correct partial commit

* rm unused

* ui: address review from allozaur

Replace the magic strings, regexes and number in the read_media parser
and service with named constants. Path splitting reuses
FILE_PATH_SEPARATOR_REGEX, the size header regex moves to
READ_MEDIA_SIZE_REGEX derived from PREFIX_SIZE, and
FILE_EXTENSION_SEPARATOR lands next to it in constants/code.ts.

---------

Co-authored-by: ckrafft <ckrafft@epyc>
Co-authored-by: Xuan Son Nguyen <son@huggingface.co>
Co-authored-by: Pascal <admin@serveurperso.com>
2026-08-12 12:03:32 +02:00

2.7 KiB

Agentic thread perf harness

Two tiers, both reusing the existing vitest projects (see vite.config.ts).

Tier 1 - agentic-stream.perf.svelte.test.ts (project: client, real Chromium)

Mounts ChatMessageAgenticContent and replays a stream, replacing the message object on each chunk exactly as the real pipeline does:

  • chat.svelte.ts updateStreamingUI() runs per SSE chunk
  • conversations.svelte.ts updateMessageAtIndex does { ...old, ...updates }

That new object identity is the thing under test: it cascades through deriveAgenticSections (which returns fresh AgenticSection objects) into every tool-call block in the message, including completed ones.

npx vitest --project=client --run tests/client/agentic-stream.perf.svelte.test.ts

Reading the output

  • mean / p95 / max - the synchronous window per token: prop write, await tick(), then a forced offsetHeight read so style and layout are included rather than deferred.
  • sync - sum of those windows. This is the number to optimize.
  • wall - the whole run including work MarkdownContent defers into its own requestAnimationFrame. It carries a ~16.7ms/token idle floor because the harness yields a frame each iteration, so compare wall across fixtures, never against sync.

The knobs, and what each one discriminates

The point of the harness is the scaling curve, not any single number.

Knob Reads on
priorToolCalls (0/1/5/20) the reactive fan-out. Flat => no fan-out. Linear => confirmed.
toolResultBytes whole-blob string scans (extractSearchResults, parseToolResultWithMedia, classifyToolResult).
editFileEdits computeLineDiff, the O(m*n) LCS.
openCodeFence hljs.highlightAuto on partial code.

Deliberately no hard assertions: CI timing is noisy and the value here is the before/after delta, not a gate.

Caveat

This measures one message's subtree. In the real app ChatMessages.svelte rebuilds its whole displayMessages list per token, so multiply by the number of rendered messages to get the conversation-level cost.

Tier 2 - ../unit/agentic-hotpath.bench.ts (project: unit, node)

Per-call costs for the pure functions the curve implicates.

npx vitest bench --project=unit --run tests/unit/agentic-hotpath.bench.ts