mirror of
https://github.com/ggml-org/llama.cpp.git
synced 2026-08-30 17:11:19 +02:00
4dd127584b
* server: add read_image tool (#25875) Adds a server-tool that allows vision models to analyze server-side images. This tool is reading a single file for now: The image data is base64 encoded and passed to the UI, which decodes it, fills the <img> tag and removes the data URI before passing the tool result back to the model. * cleanup read_image tool: move magic strings to constants * Add dedicated constants file: tools/ui/src/lib/constants/read-image.ts with PREFIX_IMAGE, PREFIX_SIZE, PREFIX_MIME constants * Use ATTACHMENT_SAVED_REGEX from agentic.ts in ChatMessageToolCallBlockReadImage.svelte * Use NEWLINE constant from code.ts instead of hardcoded '\n' * Use PREFIX_SIZE in regex pattern for size parsing * Add SERVER_TOOL_READ_IMAGE_PREFIX_* constants in C++ server-tools.cpp to match the TypeScript PREFIX_* constants for consistency * server: rename read_image tool to read_media for images and audio * Rename server_tool_read_image to server_tool_read_media in C++ * Rename enum BuiltInTool.READ_IMAGE to READ_MEDIA * Rename UI constants, parser, and Svelte component files * Update display label from 'Read image' to 'Read media' * ui: consolidate audio data URI handling into shared utility * Extract getAudioInputFormat to a shared utility (was duplicated inline) * Store raw base64 in base64Data on the message object * Use base64Data to construct data URIs for audio rendering * Update agentic store to build INPUT_AUDIO parts from base64Data * server: read_media: restrict audio to wav/mp3 and minor fixes * Server get_mime_from_extension now only advertises audio/wav and audio/mpeg (the only formats the model's input_audio API accepts) * Case-insensitive extension matching (fixes .MP3, .Wav, etc.) * Unknown extensions return an error instead of a multi-MB data URI that inflates model context with garbage * Updated tool description to document supported formats * Frontend AUDIO_MIME_TO_EXTENSION trimmed to match server * fix a missing import in tools/ui/src/lib/stores/agentic.svelte.ts * server: read_media: add to --tools help text and README tool list * ui: fix indentation in ChatMessageToolCallBlockDefault.svelte * server: read_media tool: fix a cast to use the correct type * server: read_media: multiple fixes * server-tools.cpp import cctype, remove UTF-8 char, check mime before reading file * ui: add MimeTypePrefix.AUDIO and use it in agentic.svelte.ts * server: make read_media inherit from read_file and add uses_cwd * ui: fix formating issues * rm from server * move it to frontend-only tool * correct partial commit * rm unused * ui: address review from allozaur Replace the magic strings, regexes and number in the read_media parser and service with named constants. Path splitting reuses FILE_PATH_SEPARATOR_REGEX, the size header regex moves to READ_MEDIA_SIZE_REGEX derived from PREFIX_SIZE, and FILE_EXTENSION_SEPARATOR lands next to it in constants/code.ts. --------- Co-authored-by: ckrafft <ckrafft@epyc> Co-authored-by: Xuan Son Nguyen <son@huggingface.co> Co-authored-by: Pascal <admin@serveurperso.com>
59 lines
2.7 KiB
Markdown
59 lines
2.7 KiB
Markdown
# Agentic thread perf harness
|
|
|
|
Two tiers, both reusing the existing vitest projects (see `vite.config.ts`).
|
|
|
|
## Tier 1 - `agentic-stream.perf.svelte.test.ts` (project: `client`, real Chromium)
|
|
|
|
Mounts `ChatMessageAgenticContent` and replays a stream, replacing the message
|
|
object on each chunk exactly as the real pipeline does:
|
|
|
|
- `chat.svelte.ts` `updateStreamingUI()` runs per SSE chunk
|
|
- `conversations.svelte.ts` `updateMessageAtIndex` does `{ ...old, ...updates }`
|
|
|
|
That new object identity is the thing under test: it cascades through
|
|
`deriveAgenticSections` (which returns fresh `AgenticSection` objects) into every
|
|
tool-call block in the message, including completed ones.
|
|
|
|
```
|
|
npx vitest --project=client --run tests/client/agentic-stream.perf.svelte.test.ts
|
|
```
|
|
|
|
### Reading the output
|
|
|
|
- `mean` / `p95` / `max` - the synchronous window per token: prop write,
|
|
`await tick()`, then a forced `offsetHeight` read so style and layout are
|
|
included rather than deferred.
|
|
- `sync` - sum of those windows. This is the number to optimize.
|
|
- `wall` - the whole run including work `MarkdownContent` defers into its own
|
|
`requestAnimationFrame`. It carries a ~16.7ms/token idle floor because the
|
|
harness yields a frame each iteration, so compare `wall` **across fixtures**,
|
|
never against `sync`.
|
|
|
|
### The knobs, and what each one discriminates
|
|
|
|
The point of the harness is the _scaling curve_, not any single number.
|
|
|
|
| Knob | Reads on |
|
|
| --------------------------- | --------------------------------------------------------------------------------------------------- |
|
|
| `priorToolCalls` (0/1/5/20) | the reactive fan-out. Flat => no fan-out. Linear => confirmed. |
|
|
| `toolResultBytes` | whole-blob string scans (`extractSearchResults`, `parseToolResultWithMedia`, `classifyToolResult`). |
|
|
| `editFileEdits` | `computeLineDiff`, the O(m\*n) LCS. |
|
|
| `openCodeFence` | `hljs.highlightAuto` on partial code. |
|
|
|
|
Deliberately no hard assertions: CI timing is noisy and the value here is the
|
|
before/after delta, not a gate.
|
|
|
|
### Caveat
|
|
|
|
This measures one message's subtree. In the real app `ChatMessages.svelte`
|
|
rebuilds its whole `displayMessages` list per token, so multiply by the number
|
|
of rendered messages to get the conversation-level cost.
|
|
|
|
## Tier 2 - `../unit/agentic-hotpath.bench.ts` (project: `unit`, node)
|
|
|
|
Per-call costs for the pure functions the curve implicates.
|
|
|
|
```
|
|
npx vitest bench --project=unit --run tests/unit/agentic-hotpath.bench.ts
|
|
```
|