Files
llama.cpp/tools/tuning/bench.h
T
YiChen Lv f280b26983 metal : per-device tuned (Q, NE) for flash-attn vec (#26570)
* metal : per-device tuned (Q, NE) for flash-attn vec (#25750)

* rebase Q-generic FA vec body from 01dc93607 (#23114)

* add 53 f16 (Q,NE) flash-attn vec instantiations (vec 80 -> 133)

* add FA vec (Q,NE) tuning table + dispatch wiring + SMEM cap fallback

* add  FA vec (Q,NE) perf sweep

* fill tuning result

* fold family table into a per-family representative SKU

* refactor tuning result format

* extend FA vec tuning to quantized KV caches

* sync fa vec tuner bucketing with runtime, use pointwise tuning regret

* update tuned table

* format and cleanup

* prefix fa_vec tuning procs with ggml_backend_metal_tuning_, drop unused fa_vec_override_active

* add device id -> token lookup for the offline tuning tool

* add ggml-metal-tuning skeleton

* add op-agnostic perf cell + median timing for the tuner

* add FA-vec graph build + tensor init to the tuner

* tools : add FA-vec (Q,NE) sweep, compression and table emit

* cool down and re-measure the dirty window on thermal drift

* test-backend-ops : replace the FA vec tune mode with a bounded (Q,NE) slice

* tools : document the Metal tuner, point the table comment at it

* abort on unknown KV type, single-source fa_vec_legal_ne

* cleanup

* honor -o in the FA vec (Q,NE) slice

* retune FA-vec (Q, NE) under a pointwise no-harm gate

* cont : add fa-vec tunings for M1 Pro, M2 Ultra, M5 Max

---------

Co-authored-by: Georgi Gerganov <ggerganov@gmail.com>
2026-08-24 19:22:27 +03:00

58 lines
2.1 KiB
C++

#pragma once
#include "ggml-backend.h"
#include "ggml-cpp.h"
#include "ggml.h"
#include <cstdint>
#include <functional>
#include <vector>
// A prebuilt graph replicated to amortize dispatch and synchronization overhead.
struct perf_cell {
ggml_context_ptr ctx;
ggml_backend_buffer_ptr buf;
ggml_cgraph * gf = nullptr;
int n_runs = 0;
};
using build_graph_fn = std::function<ggml_tensor *(ggml_context *)>;
using init_tensors_fn = std::function<void(ggml_context *)>;
using op_flops_fn = std::function<uint64_t(ggml_tensor *)>;
perf_cell build_perf_cell(ggml_backend_t backend,
const build_graph_fn & build,
const init_tensors_fn & init,
const op_flops_fn & flops);
double time_cell_median(ggml_backend_t backend, const perf_cell & cell, int reps);
struct cooldown_opts {
bool enabled = true;
double drift = 0.10; // anchor drift that triggers a cooldown
double eps = 0.03; // anchor tolerance to call the GPU cool again
int max_wait = 120; // seconds of cooling per cell before giving up
int max_retry = 2; // re-measure rounds per cell before giving up
};
using set_candidate_fn = std::function<void(int)>;
using clear_candidate_fn = std::function<void()>;
struct cell_result {
std::vector<double> t;
bool trusted = true;
double anchor_min = 0.0;
double anchor_max = 0.0;
};
// Times candidates in order while using baseline_cand as a thermal-drift anchor.
cell_result measure_cell(ggml_backend_t backend,
const perf_cell & cell,
int reps,
const std::vector<int> & order,
const set_candidate_fn & set_cand,
const clear_candidate_fn & clear_cand,
int baseline_cand,
const cooldown_opts & cool,
const char * cell_label);