* server: add read_image tool (#25875) Adds a server-tool that allows vision models to analyze server-side images. This tool is reading a single file for now: The image data is base64 encoded and passed to the UI, which decodes it, fills the <img> tag and removes the data URI before passing the tool result back to the model. * cleanup read_image tool: move magic strings to constants * Add dedicated constants file: tools/ui/src/lib/constants/read-image.ts with PREFIX_IMAGE, PREFIX_SIZE, PREFIX_MIME constants * Use ATTACHMENT_SAVED_REGEX from agentic.ts in ChatMessageToolCallBlockReadImage.svelte * Use NEWLINE constant from code.ts instead of hardcoded '\n' * Use PREFIX_SIZE in regex pattern for size parsing * Add SERVER_TOOL_READ_IMAGE_PREFIX_* constants in C++ server-tools.cpp to match the TypeScript PREFIX_* constants for consistency * server: rename read_image tool to read_media for images and audio * Rename server_tool_read_image to server_tool_read_media in C++ * Rename enum BuiltInTool.READ_IMAGE to READ_MEDIA * Rename UI constants, parser, and Svelte component files * Update display label from 'Read image' to 'Read media' * ui: consolidate audio data URI handling into shared utility * Extract getAudioInputFormat to a shared utility (was duplicated inline) * Store raw base64 in base64Data on the message object * Use base64Data to construct data URIs for audio rendering * Update agentic store to build INPUT_AUDIO parts from base64Data * server: read_media: restrict audio to wav/mp3 and minor fixes * Server get_mime_from_extension now only advertises audio/wav and audio/mpeg (the only formats the model's input_audio API accepts) * Case-insensitive extension matching (fixes .MP3, .Wav, etc.) * Unknown extensions return an error instead of a multi-MB data URI that inflates model context with garbage * Updated tool description to document supported formats * Frontend AUDIO_MIME_TO_EXTENSION trimmed to match server * fix a missing import in tools/ui/src/lib/stores/agentic.svelte.ts * server: read_media: add to --tools help text and README tool list * ui: fix indentation in ChatMessageToolCallBlockDefault.svelte * server: read_media tool: fix a cast to use the correct type * server: read_media: multiple fixes * server-tools.cpp import cctype, remove UTF-8 char, check mime before reading file * ui: add MimeTypePrefix.AUDIO and use it in agentic.svelte.ts * server: make read_media inherit from read_file and add uses_cwd * ui: fix formating issues * rm from server * move it to frontend-only tool * correct partial commit * rm unused * ui: address review from allozaur Replace the magic strings, regexes and number in the read_media parser and service with named constants. Path splitting reuses FILE_PATH_SEPARATOR_REGEX, the size header regex moves to READ_MEDIA_SIZE_REGEX derived from PREFIX_SIZE, and FILE_EXTENSION_SEPARATOR lands next to it in constants/code.ts. --------- Co-authored-by: ckrafft <ckrafft@epyc> Co-authored-by: Xuan Son Nguyen <son@huggingface.co> Co-authored-by: Pascal <admin@serveurperso.com>
2.7 KiB
Agentic thread perf harness
Two tiers, both reusing the existing vitest projects (see vite.config.ts).
Tier 1 - agentic-stream.perf.svelte.test.ts (project: client, real Chromium)
Mounts ChatMessageAgenticContent and replays a stream, replacing the message
object on each chunk exactly as the real pipeline does:
chat.svelte.tsupdateStreamingUI()runs per SSE chunkconversations.svelte.tsupdateMessageAtIndexdoes{ ...old, ...updates }
That new object identity is the thing under test: it cascades through
deriveAgenticSections (which returns fresh AgenticSection objects) into every
tool-call block in the message, including completed ones.
npx vitest --project=client --run tests/client/agentic-stream.perf.svelte.test.ts
Reading the output
mean/p95/max- the synchronous window per token: prop write,await tick(), then a forcedoffsetHeightread so style and layout are included rather than deferred.sync- sum of those windows. This is the number to optimize.wall- the whole run including workMarkdownContentdefers into its ownrequestAnimationFrame. It carries a ~16.7ms/token idle floor because the harness yields a frame each iteration, so comparewallacross fixtures, never againstsync.
The knobs, and what each one discriminates
The point of the harness is the scaling curve, not any single number.
| Knob | Reads on |
|---|---|
priorToolCalls (0/1/5/20) |
the reactive fan-out. Flat => no fan-out. Linear => confirmed. |
toolResultBytes |
whole-blob string scans (extractSearchResults, parseToolResultWithMedia, classifyToolResult). |
editFileEdits |
computeLineDiff, the O(m*n) LCS. |
openCodeFence |
hljs.highlightAuto on partial code. |
Deliberately no hard assertions: CI timing is noisy and the value here is the before/after delta, not a gate.
Caveat
This measures one message's subtree. In the real app ChatMessages.svelte
rebuilds its whole displayMessages list per token, so multiply by the number
of rendered messages to get the conversation-level cost.
Tier 2 - ../unit/agentic-hotpath.bench.ts (project: unit, node)
Per-call costs for the pure functions the curve implicates.
npx vitest bench --project=unit --run tests/unit/agentic-hotpath.bench.ts