Files
llama.cpp/tools/ui/tests/client/README-perf.md
T
parabelboi 4dd127584b ui: add read_media tool (#25877)
* server: add read_image tool (#25875)

Adds a server-tool that allows vision models to analyze server-side images.
This tool is reading a single file for now:
The image data is base64 encoded and passed to the UI, which
decodes it, fills the <img> tag and removes the data URI before
passing the tool result back to the model.

* cleanup read_image tool: move magic strings to constants

* Add dedicated constants file: tools/ui/src/lib/constants/read-image.ts
  with PREFIX_IMAGE, PREFIX_SIZE, PREFIX_MIME constants
* Use ATTACHMENT_SAVED_REGEX from agentic.ts in ChatMessageToolCallBlockReadImage.svelte
* Use NEWLINE constant from code.ts instead of hardcoded '\n'
* Use PREFIX_SIZE in regex pattern for size parsing
* Add SERVER_TOOL_READ_IMAGE_PREFIX_* constants in C++ server-tools.cpp
  to match the TypeScript PREFIX_* constants for consistency

* server: rename read_image tool to read_media for images and audio

* Rename server_tool_read_image to server_tool_read_media in C++
* Rename enum BuiltInTool.READ_IMAGE to READ_MEDIA
* Rename UI constants, parser, and Svelte component files
* Update display label from 'Read image' to 'Read media'

* ui: consolidate audio data URI handling into shared utility

* Extract getAudioInputFormat to a shared utility (was duplicated inline)
* Store raw base64 in base64Data on the message object
* Use base64Data to construct data URIs for audio rendering
* Update agentic store to build INPUT_AUDIO parts from base64Data

* server: read_media: restrict audio to wav/mp3 and minor fixes

* Server get_mime_from_extension now only advertises audio/wav and
  audio/mpeg (the only formats the model's input_audio API accepts)
* Case-insensitive extension matching (fixes .MP3, .Wav, etc.)
* Unknown extensions return an error instead of a multi-MB data URI
  that inflates model context with garbage
* Updated tool description to document supported formats
* Frontend AUDIO_MIME_TO_EXTENSION trimmed to match server
* fix a missing import in tools/ui/src/lib/stores/agentic.svelte.ts

* server: read_media: add to --tools help text and README tool list

* ui: fix indentation in ChatMessageToolCallBlockDefault.svelte

* server: read_media tool: fix a cast to use the correct type

* server: read_media: multiple fixes

* server-tools.cpp import cctype, remove UTF-8 char, check mime before reading file
* ui: add MimeTypePrefix.AUDIO and use it in agentic.svelte.ts

* server: make read_media inherit from read_file and add uses_cwd

* ui: fix formating issues

* rm from server

* move it to frontend-only tool

* correct partial commit

* rm unused

* ui: address review from allozaur

Replace the magic strings, regexes and number in the read_media parser
and service with named constants. Path splitting reuses
FILE_PATH_SEPARATOR_REGEX, the size header regex moves to
READ_MEDIA_SIZE_REGEX derived from PREFIX_SIZE, and
FILE_EXTENSION_SEPARATOR lands next to it in constants/code.ts.

---------

Co-authored-by: ckrafft <ckrafft@epyc>
Co-authored-by: Xuan Son Nguyen <son@huggingface.co>
Co-authored-by: Pascal <admin@serveurperso.com>
2026-08-12 12:03:32 +02:00

59 lines
2.7 KiB
Markdown

# Agentic thread perf harness
Two tiers, both reusing the existing vitest projects (see `vite.config.ts`).
## Tier 1 - `agentic-stream.perf.svelte.test.ts` (project: `client`, real Chromium)
Mounts `ChatMessageAgenticContent` and replays a stream, replacing the message
object on each chunk exactly as the real pipeline does:
- `chat.svelte.ts` `updateStreamingUI()` runs per SSE chunk
- `conversations.svelte.ts` `updateMessageAtIndex` does `{ ...old, ...updates }`
That new object identity is the thing under test: it cascades through
`deriveAgenticSections` (which returns fresh `AgenticSection` objects) into every
tool-call block in the message, including completed ones.
```
npx vitest --project=client --run tests/client/agentic-stream.perf.svelte.test.ts
```
### Reading the output
- `mean` / `p95` / `max` - the synchronous window per token: prop write,
`await tick()`, then a forced `offsetHeight` read so style and layout are
included rather than deferred.
- `sync` - sum of those windows. This is the number to optimize.
- `wall` - the whole run including work `MarkdownContent` defers into its own
`requestAnimationFrame`. It carries a ~16.7ms/token idle floor because the
harness yields a frame each iteration, so compare `wall` **across fixtures**,
never against `sync`.
### The knobs, and what each one discriminates
The point of the harness is the _scaling curve_, not any single number.
| Knob | Reads on |
| --------------------------- | --------------------------------------------------------------------------------------------------- |
| `priorToolCalls` (0/1/5/20) | the reactive fan-out. Flat => no fan-out. Linear => confirmed. |
| `toolResultBytes` | whole-blob string scans (`extractSearchResults`, `parseToolResultWithMedia`, `classifyToolResult`). |
| `editFileEdits` | `computeLineDiff`, the O(m\*n) LCS. |
| `openCodeFence` | `hljs.highlightAuto` on partial code. |
Deliberately no hard assertions: CI timing is noisy and the value here is the
before/after delta, not a gate.
### Caveat
This measures one message's subtree. In the real app `ChatMessages.svelte`
rebuilds its whole `displayMessages` list per token, so multiply by the number
of rendered messages to get the conversation-level cost.
## Tier 2 - `../unit/agentic-hotpath.bench.ts` (project: `unit`, node)
Per-call costs for the pure functions the curve implicates.
```
npx vitest bench --project=unit --run tests/unit/agentic-hotpath.bench.ts
```