mirror of
https://github.com/LostRuins/koboldcpp.git
synced 2026-09-04 20:11:17 +02:00
62bf73d25c
* Get started with Onyx * Add architecture * Skip keys handled in super() * Loading tensors * Shorten * Graph * Apply suggestion from @pcuenca * Remove norm now embedding in transformers weights * Add eot * Explicit output_multiplier * Handle post_norm_eps * No super call; unhardcode eot. The pattern `self._set_vocab_gpt2()` seems preferred throughout the codebase, and it allows `set_vocab()` to be called from a different part of the Python class hierarchy: the drafter model converter that we may need eventually. * Register for drafting * DFlash: inherit rope type from the linked target. Another option would be to store it in the gguf file itself. * mmproj conversion Note: some fields to be renamed after the implementation works. We are keeping compatibility with the reference Meta gguf for testing purposes. * "clip" header declarations * Load mmproj * Pre-processing * Graph * Go back to using delimiters. Otherwise our generations are worse. Transformers does not use them. We need to trace inputs to verify whether they are equivalent. * downsample_factor -> merge_size * Add vision graph lol, forgot from a previous commit * Additional renames, align with llama.cpp / transformers * Prefer _size instead of independent _h and _w * Fix token layout Co-authored-by: Young Han <younghan@fb.com> * onyx: bring the chat parser onto the onyx branch common/chat.cpp on this branch has no Onyx handling, so a converted model serves malformed chat: the assistant preamble leaks into content ("to=self<|message|>...") and tool calls fail with HTTP 500 "The model produced output that does not match the expected peg-native format" common_chat_params_init_onyx exists on onyx-fair-patch, added there by 8bb73dd3d. It was never on this branch, so this is not a regression -- the two lines developed independently. The code here is taken verbatim from that commit. It is the clean side of `git merge origin/onyx-fair-patch`: chat.cpp is one of the files that merges without conflict. The full merge is not viable -- it produces 13 conflicts, including add/add on conversion/onyx.py and src/models/onyx.cpp where the q_norm-folding and metadata-scale approaches contradict each other, and #4/#7 are stacked on this branch's side of that. Verified on this branch: builds with 0 errors, converts an Onyx checkpoint, and serving it gives "4" for "What is 2+2?" plus a correct get_weather {"city":"Paris"} tool call, where the unported branch gives the two failures above. No converter or runtime changes are included, so this should not interact with the q_norm work. Co-authored-by: Beto de Paola <betodepaola@meta.com> * Less params, bilinear pos-emb interpolation as a graph op instead of CPU * Map to symbolic V_MMPROJ instead of strings * Make a couple params explicit * Patchify via build_inp() * No param for rope_theta * Small cleanup * Restore blank line * Unpermute, to adapt to the latest transformers checkpoint * Apply norm after token embeddings This follows the latest transformers approach. * Remove duplicated function * build_vit * onyx: use the model rope theta on sliding-window layers * DFlash: conversion from transformers drafter * Revert rope_type derivation from target NOTE: this breaks compatibility with Meta's distributed DFlash GGUFs, as the Q/K are stored in "NEOX" (rotated half) format, like in transformers. * Apply suggestion from @pcuenca * Set model type * Remove comment that will become obsolete * Hardcode post_norm_rms_eps instead of new param * Derive SWA+RoPE pattern from gguf array or scalar * Fix model type <-> number of layers * Reorder * Rename * Fix typo * DFlash: seed the draft KV cache from multimodal embedding batches `common_speculative_impl_draft_dflash::process()` returned early on any batch carrying embeddings, so an image prefill never had its target-layer features fused through the DFlash encoder and injected into the draft's KV cache. That left a hole spanning the image's positions, and the next injection at a post-image position failed to initialize its batch: ``` decoding image batch 1/1, n_tokens_batch = 256 decode: failed to initialize batch llama_decode: failed to decode, ret = -1 process: llama_decode(ctx_dft) failed rc=-1 (n_tokens=17, offset=0) srv decode: failed to process speculative batch ``` Every image request with `--spec-type draft-dflash` failed with HTTP 500. Text-only was unaffected, since those batches carry token ids and were let through. Restore the earlier condition, which admits a batch that is either tokens or embeddings and skips only the degenerate neither/both cases. The rest of `process()` is already layout-agnostic -- it gathers features via `llama_get_embeddings_layer_inp()` and indexes `batch_in.pos[]` / `batch_in.seq_id[]`, none of which assume token ids -- so this is the whole fix. Validated against `muse-glimmer-30B-bf16.gguf` + `mmproj-muse-glimmer-30B-bf16.gguf` + a DFlash draft head, on an image describe-the-shapes request: - before: HTTP 500, `failed to process speculative batch` - after: HTTP 200, draft acceptance 0.34012 (167 accepted / 491 generated), mean len 3.04 Output equivalence holds, which is the property that matters: at temperature 0 the drafted response is byte-identical to the same request served with no draft attached (1213/1213 chars), so the draft is drafting correctly through the image context rather than merely not crashing. * Conversion: prefer rewrite to mapping * Revert "Conversion: prefer rewrite to mapping" This reverts commit a92d0ac584d315e876741e85b6dad3dbc8b23bf7. * fix lint * sliding_window metadata is not optional * disable state save/load * Apply suggestion from @pcuenca --------- Co-authored-by: Young Han <younghan@fb.com> Co-authored-by: Beto de Paola <betodepaola@meta.com> Co-authored-by: Daniel Han <michaelhan2050@gmail.com> Co-authored-by: ruanrms <ruanslv@gmail.com> Co-authored-by: Xuan Son Nguyen <son@huggingface.co> Co-authored-by: Sigbjørn Skjæret <sigbjorn.skjaeret@huggingface.co>
239 lines
11 KiB
C++
239 lines
11 KiB
C++
#pragma once
|
|
|
|
#include "ggml.h"
|
|
#include "clip-model.h"
|
|
|
|
#include <vector>
|
|
#include <string>
|
|
|
|
#define MTMD_INTERNAL_HEADER
|
|
|
|
struct mtmd_image_preproc_out {
|
|
std::vector<clip_image_f32> entries;
|
|
// grid size is required for llava-uhd style models
|
|
|
|
clip_image_f32 overview; // overview image (downscaled image)
|
|
int grid_x = 0;
|
|
int grid_y = 0;
|
|
|
|
void append(const clip_hparams & hparams, const clip_image_u8 & img, bool normalized = true);
|
|
void append(const clip_hparams & hparams, const std::vector<clip_image_u8> & imgs, bool normalized = true);
|
|
void append(const clip_hparams & hparams, clip_image_f32 & img, bool normalized = true);
|
|
|
|
void append_overview(const clip_hparams & hparams, const clip_image_u8 & img, bool normalized = true);
|
|
bool has_overview() const {
|
|
return overview.nx() > 0 || overview.ny() > 0;
|
|
}
|
|
};
|
|
|
|
// base class, models must inherit from this class
|
|
struct mtmd_image_preprocessor {
|
|
const clip_hparams & hparams;
|
|
|
|
mtmd_image_preprocessor(const clip_ctx * ctx): hparams(*clip_get_hparams(ctx)) {}
|
|
|
|
virtual ~mtmd_image_preprocessor() = default;
|
|
virtual mtmd_image_preproc_out preprocess(const clip_image_u8 & img) = 0;
|
|
};
|
|
|
|
/**
|
|
* implementation of LLaVA-UHD:
|
|
* - https://arxiv.org/pdf/2403.11703
|
|
* - https://github.com/thunlp/LLaVA-UHD
|
|
* - https://github.com/thunlp/LLaVA-UHD/blob/302301bc2175f7e717fb8548516188e89f649753/llava_uhd/train/llava-uhd/slice_logic.py#L118
|
|
*
|
|
* overview:
|
|
* - an image always have a single overview (downscaled image)
|
|
* - an image can have 0 or multiple slices, depending on the image size
|
|
* - each slice can then be considered as a separate image
|
|
*
|
|
* note: the term "slice" and "tile" are used interchangeably
|
|
*
|
|
* for example:
|
|
*
|
|
* [overview] --> [slice 1] --> [slice 2]
|
|
* | |
|
|
* +--> [slice 3] --> [slice 4]
|
|
*
|
|
* NOTE: for the ordering of overview, set "ov_img_first" on the mtmd_context
|
|
*/
|
|
struct mtmd_image_preprocessor_llava_uhd : mtmd_image_preprocessor {
|
|
mtmd_image_preprocessor_llava_uhd(const clip_ctx * ctx) : mtmd_image_preprocessor(ctx) {}
|
|
mtmd_image_preproc_out preprocess(const clip_image_u8 & img) override;
|
|
|
|
struct slice_coordinates {
|
|
int x;
|
|
int y;
|
|
clip_image_size size;
|
|
};
|
|
|
|
struct slice_instructions {
|
|
clip_image_size overview_size; // size of downscaled image
|
|
clip_image_size refined_size; // size of image right before slicing (must be multiple of slice size)
|
|
clip_image_size grid_size; // grid_size.width * grid_size.height = number of slices
|
|
std::vector<slice_coordinates> slices;
|
|
};
|
|
|
|
virtual slice_instructions get_slice_instructions(const clip_image_size & original_size);
|
|
|
|
struct slice_output {
|
|
clip_image_u8 overview;
|
|
std::vector<clip_image_u8> slices;
|
|
};
|
|
slice_output slice_image(const clip_image_u8 & img, const slice_instructions & inst);
|
|
|
|
protected:
|
|
clip_image_size get_best_resize(const clip_image_size & original_size, int scale_resolution, int patch_size, bool allow_upscale = false);
|
|
|
|
private:
|
|
clip_image_size resize_maintain_aspect_ratio(const clip_image_size & orig, const clip_image_size & target_max);
|
|
|
|
/**
|
|
* Selects the best resolution from a list of possible resolutions based on the original size.
|
|
*
|
|
* For example, when given a list of resolutions:
|
|
* - 100x100
|
|
* - 200x100
|
|
* - 100x200
|
|
* - 200x200
|
|
*
|
|
* And an input image of size 111x200, then 100x200 is the best fit (least wasted resolution).
|
|
*
|
|
* @param original_size The original size of the image
|
|
* @param possible_resolutions A list of possible resolutions
|
|
* @return The best fit resolution
|
|
*/
|
|
clip_image_size select_best_resolution(const clip_image_size & original_size, const std::vector<clip_image_size> & possible_resolutions);
|
|
int ensure_divide(int length, int patch_size);
|
|
clip_image_size get_refine_size(const clip_image_size & original_size, const clip_image_size & grid, int scale_resolution, int patch_size, bool allow_upscale = false);
|
|
clip_image_size get_best_grid(const int max_slice_nums, const int multiple, const float log_ratio);
|
|
};
|
|
|
|
// downscale or upscale the input image to fixed size
|
|
struct mtmd_image_preprocessor_fixed_size : mtmd_image_preprocessor {
|
|
mtmd_image_preprocessor_fixed_size(const clip_ctx * ctx) : mtmd_image_preprocessor(ctx) {}
|
|
mtmd_image_preproc_out preprocess(const clip_image_u8 & img) override;
|
|
};
|
|
|
|
// resize image to multiple of patch_size*n_merge, while preserving aspect ratio
|
|
// if image_resize_pad is true, the resized image will be padded, otherwise it will be either stretched or center-cropped depending on image_resize_pad
|
|
// this is used by models with native support for dynamic image size, for example: Qwen-VL, Pixtral, Kimi-VL, etc
|
|
struct mtmd_image_preprocessor_dyn_size : mtmd_image_preprocessor {
|
|
mtmd_image_preprocessor_dyn_size(const clip_ctx * ctx) : mtmd_image_preprocessor(ctx) {}
|
|
mtmd_image_preproc_out preprocess(const clip_image_u8 & img) override;
|
|
};
|
|
|
|
// similar to mtmd_image_preprocessor_dyn_size, but resize the image to have longest edge equal to hparams.image_longest_edge, while preserving aspect ratio
|
|
struct mtmd_image_preprocessor_longest_edge : mtmd_image_preprocessor {
|
|
mtmd_image_preprocessor_longest_edge(const clip_ctx * ctx) : mtmd_image_preprocessor(ctx) {}
|
|
mtmd_image_preproc_out preprocess(const clip_image_u8 & img) override;
|
|
};
|
|
|
|
// custom llava-uhd slicing logic for MiniCPM-V
|
|
struct mtmd_image_preprocessor_minicpmv : mtmd_image_preprocessor_llava_uhd {
|
|
using mtmd_image_preprocessor_llava_uhd::mtmd_image_preprocessor_llava_uhd;
|
|
slice_instructions get_slice_instructions(const clip_image_size & original_size) override;
|
|
};
|
|
|
|
// custom llava-uhd slicing logic for LFM2
|
|
// ref: https://github.com/huggingface/transformers/blob/v5.1.0/src/transformers/models/lfm2_vl/image_processing_lfm2_vl_fast.py
|
|
struct mtmd_image_preprocessor_lfm2 : mtmd_image_preprocessor_llava_uhd {
|
|
// ref: https://huggingface.co/LiquidAI/LFM2.5-VL-1.6B/blob/main/processor_config.json
|
|
static constexpr int min_tiles = 2;
|
|
static constexpr int max_tiles = 10;
|
|
static constexpr float max_pixels_tolerance = 2.0f;
|
|
static constexpr int tile_size = 512;
|
|
|
|
using mtmd_image_preprocessor_llava_uhd::mtmd_image_preprocessor_llava_uhd;
|
|
slice_instructions get_slice_instructions(const clip_image_size & original_size) override;
|
|
|
|
private:
|
|
clip_image_size find_closest_aspect_ratio(
|
|
float aspect_ratio,
|
|
const std::vector<clip_image_size> & target_ratios,
|
|
int width, int height);
|
|
std::vector<clip_image_size> get_target_ratios();
|
|
clip_image_size get_grid_layout(int height, int width);
|
|
};
|
|
|
|
struct mtmd_image_preprocessor_idefics3 : mtmd_image_preprocessor_llava_uhd {
|
|
mtmd_image_preprocessor_idefics3(const clip_ctx * ctx) : mtmd_image_preprocessor_llava_uhd(ctx) {}
|
|
mtmd_image_preproc_out preprocess(const clip_image_u8 & img) override;
|
|
};
|
|
|
|
struct mtmd_image_preprocessor_internvl : mtmd_image_preprocessor_llava_uhd {
|
|
mtmd_image_preprocessor_internvl(const clip_ctx * ctx) : mtmd_image_preprocessor_llava_uhd(ctx) {}
|
|
mtmd_image_preproc_out preprocess(const clip_image_u8 & img) override;
|
|
};
|
|
|
|
// DeepSeek-OCR (v1/v2) global view + optional local tile grid
|
|
struct mtmd_image_preprocessor_deepseekocr : mtmd_image_preprocessor {
|
|
mtmd_image_preprocessor_deepseekocr(const clip_ctx * ctx)
|
|
: mtmd_image_preprocessor(ctx),
|
|
fuse_row(clip_get_projector_type(ctx) == PROJECTOR_TYPE_DEEPSEEKOCR),
|
|
base_size(hparams.image_size),
|
|
tile_size(hparams.preproc_tile_size),
|
|
min_tiles(hparams.preproc_min_tiles),
|
|
max_tiles(hparams.preproc_max_tiles) {}
|
|
mtmd_image_preproc_out preprocess(const clip_image_u8 & img) override;
|
|
|
|
private:
|
|
bool fuse_row; // v1 fuses a tile-row into one image; v2 keeps tiles separate
|
|
int base_size; // global view
|
|
int tile_size; // each tile
|
|
int min_tiles;
|
|
int max_tiles;
|
|
|
|
std::vector<clip_image_size> get_target_ratios() const;
|
|
clip_image_size find_closest_aspect_ratio(
|
|
float aspect_ratio,
|
|
const std::vector<clip_image_size> & target_ratios,
|
|
int width, int height) const;
|
|
};
|
|
|
|
// custom image preprocessing for Step3VL
|
|
// ref: https://huggingface.co/stepfun-ai/Step3-VL-10B/blob/main/processing_step3.py
|
|
struct mtmd_image_preprocessor_step3vl : mtmd_image_preprocessor_llava_uhd {
|
|
mtmd_image_preprocessor_step3vl(const clip_ctx * ctx) : mtmd_image_preprocessor_llava_uhd(ctx) {}
|
|
mtmd_image_preproc_out preprocess(const clip_image_u8 & img) override;
|
|
static slice_instructions build_slice_instructions(const clip_hparams & params, const clip_image_size & prepared_size);
|
|
|
|
private:
|
|
static constexpr int default_image_longest_edge = 3024;
|
|
static constexpr int default_image_crop_size = 504;
|
|
static constexpr float small_aspect_ratio_limit = 1.5f;
|
|
static constexpr float wide_aspect_ratio_limit = 4.0f;
|
|
static constexpr float crop_rounding_threshold = 0.2f;
|
|
|
|
void img_u8_resize_bilinear_to_f32(
|
|
const clip_image_u8 & src,
|
|
clip_image_f32 & dst,
|
|
int target_width,
|
|
int target_height,
|
|
const float mean[3],
|
|
const float std[3]);
|
|
static int get_image_longest_edge(const clip_hparams & params);
|
|
static int determine_window_size(const clip_hparams & params, int longer, int shorter);
|
|
static int calc_crop_extent(int length, int window_size);
|
|
static std::vector<int> calc_grid(int length, int window_size);
|
|
static clip_image_u8 prepare_image(const clip_image_u8 & img, const clip_hparams & params);
|
|
static clip_image_u8 crop_with_black_padding(const clip_image_u8 & image, int x, int y, int w, int h);
|
|
};
|
|
|
|
struct mtmd_image_preprocessor_youtuvl : mtmd_image_preprocessor {
|
|
mtmd_image_preprocessor_youtuvl(const clip_ctx * ctx) : mtmd_image_preprocessor(ctx) {}
|
|
mtmd_image_preproc_out preprocess(const clip_image_u8 & img) override;
|
|
};
|
|
|
|
// similar to llava_uhd, but has add_newline
|
|
struct mtmd_image_preprocessor_granite : mtmd_image_preprocessor_llava_uhd {
|
|
mtmd_image_preprocessor_granite(const clip_ctx * ctx) : mtmd_image_preprocessor_llava_uhd(ctx) {}
|
|
mtmd_image_preproc_out preprocess(const clip_image_u8 & img) override;
|
|
};
|
|
|
|
// pick the patch grid closest to the input aspect ratio under the per-image token cap, stretch-resize.
|
|
struct mtmd_image_preprocessor_muse_glimmer : mtmd_image_preprocessor {
|
|
mtmd_image_preprocessor_muse_glimmer(const clip_ctx * ctx) : mtmd_image_preprocessor(ctx) {}
|
|
mtmd_image_preproc_out preprocess(const clip_image_u8 & img) override;
|
|
};
|