mirror of
https://github.com/ggml-org/llama.cpp.git
synced 2026-09-08 22:09:02 +02:00
62bf73d25c
* Get started with Onyx
* Add architecture
* Skip keys handled in super()
* Loading tensors
* Shorten
* Graph
* Apply suggestion from @pcuenca
* Remove norm now embedding in transformers weights
* Add eot
* Explicit output_multiplier
* Handle post_norm_eps
* No super call; unhardcode eot.
The pattern `self._set_vocab_gpt2()` seems preferred throughout the
codebase, and it allows `set_vocab()` to be called from a different part
of the Python class hierarchy: the drafter model converter that we may
need eventually.
* Register for drafting
* DFlash: inherit rope type from the linked target.
Another option would be to store it in the gguf file itself.
* mmproj conversion
Note: some fields to be renamed after the implementation works. We are
keeping compatibility with the reference Meta gguf for testing purposes.
* "clip" header declarations
* Load mmproj
* Pre-processing
* Graph
* Go back to using delimiters.
Otherwise our generations are worse.
Transformers does not use them. We need to trace inputs to verify
whether they are equivalent.
* downsample_factor -> merge_size
* Add vision graph
lol, forgot from a previous commit
* Additional renames, align with llama.cpp / transformers
* Prefer _size instead of independent _h and _w
* Fix token layout
Co-authored-by: Young Han <younghan@fb.com>
* onyx: bring the chat parser onto the onyx branch
common/chat.cpp on this branch has no Onyx handling, so a converted model
serves malformed chat: the assistant preamble leaks into content
("to=self<|message|>...") and tool calls fail with
HTTP 500 "The model produced output that does not match the expected
peg-native format"
common_chat_params_init_onyx exists on onyx-fair-patch, added there by
8bb73dd3d. It was never on this branch, so this is not a regression --
the two lines developed independently.
The code here is taken verbatim from that commit. It is the clean side of
`git merge origin/onyx-fair-patch`: chat.cpp is one of the files that
merges without conflict. The full merge is not viable -- it produces 13
conflicts, including add/add on conversion/onyx.py and src/models/onyx.cpp
where the q_norm-folding and metadata-scale approaches contradict each
other, and #4/#7 are stacked on this branch's side of that.
Verified on this branch: builds with 0 errors, converts an Onyx checkpoint,
and serving it gives "4" for "What is 2+2?" plus a correct
get_weather {"city":"Paris"} tool call, where the unported branch gives the
two failures above.
No converter or runtime changes are included, so this should not interact
with the q_norm work.
Co-authored-by: Beto de Paola <betodepaola@meta.com>
* Less params, bilinear pos-emb interpolation as a graph op instead of CPU
* Map to symbolic V_MMPROJ instead of strings
* Make a couple params explicit
* Patchify via build_inp()
* No param for rope_theta
* Small cleanup
* Restore blank line
* Unpermute, to adapt to the latest transformers checkpoint
* Apply norm after token embeddings
This follows the latest transformers approach.
* Remove duplicated function
* build_vit
* onyx: use the model rope theta on sliding-window layers
* DFlash: conversion from transformers drafter
* Revert rope_type derivation from target
NOTE: this breaks compatibility with Meta's distributed DFlash GGUFs, as
the Q/K are stored in "NEOX" (rotated half) format, like in
transformers.
* Apply suggestion from @pcuenca
* Set model type
* Remove comment that will become obsolete
* Hardcode post_norm_rms_eps instead of new param
* Derive SWA+RoPE pattern from gguf array or scalar
* Fix model type <-> number of layers
* Reorder
* Rename
* Fix typo
* DFlash: seed the draft KV cache from multimodal embedding batches
`common_speculative_impl_draft_dflash::process()` returned early on any batch carrying embeddings, so an image prefill never had its target-layer features fused through the DFlash encoder and injected into the draft's KV cache. That left a hole spanning the image's positions, and the next injection at a post-image position failed to initialize its batch:
```
decoding image batch 1/1, n_tokens_batch = 256
decode: failed to initialize batch
llama_decode: failed to decode, ret = -1
process: llama_decode(ctx_dft) failed rc=-1 (n_tokens=17, offset=0)
srv decode: failed to process speculative batch
```
Every image request with `--spec-type draft-dflash` failed with HTTP 500. Text-only was unaffected, since those batches carry token ids and were let through.
Restore the earlier condition, which admits a batch that is either tokens or embeddings and skips only the degenerate neither/both cases. The rest of `process()` is already layout-agnostic -- it gathers features via `llama_get_embeddings_layer_inp()` and indexes `batch_in.pos[]` / `batch_in.seq_id[]`, none of which assume token ids -- so this is the whole fix.
Validated against `muse-glimmer-30B-bf16.gguf` + `mmproj-muse-glimmer-30B-bf16.gguf` + a DFlash draft head, on an image describe-the-shapes request:
- before: HTTP 500, `failed to process speculative batch`
- after: HTTP 200, draft acceptance 0.34012 (167 accepted / 491 generated), mean len 3.04
Output equivalence holds, which is the property that matters: at temperature 0 the drafted response is byte-identical to the same request served with no draft attached (1213/1213 chars), so the draft is drafting correctly through the image context rather than merely not crashing.
* Conversion: prefer rewrite to mapping
* Revert "Conversion: prefer rewrite to mapping"
This reverts commit a92d0ac584.
* fix lint
* sliding_window metadata is not optional
* disable state save/load
* Apply suggestion from @pcuenca
---------
Co-authored-by: Young Han <younghan@fb.com>
Co-authored-by: Beto de Paola <betodepaola@meta.com>
Co-authored-by: Daniel Han <michaelhan2050@gmail.com>
Co-authored-by: ruanrms <ruanslv@gmail.com>
Co-authored-by: Xuan Son Nguyen <son@huggingface.co>
Co-authored-by: Sigbjørn Skjæret <sigbjorn.skjaeret@huggingface.co>
373 lines
14 KiB
C++
373 lines
14 KiB
C++
#pragma once
|
|
|
|
#include "../clip-graph.h"
|
|
|
|
#include <map>
|
|
#include <string>
|
|
#include <utility>
|
|
#include <vector>
|
|
|
|
/*
|
|
* IMPORTANT: The mtmd module does NOT accept pull requests that are fully or predominantly AI-generated.
|
|
* We encourage human contributors to ensure the quality and reliability of the codebase.
|
|
*/
|
|
|
|
struct clip_graph_siglip : clip_graph {
|
|
clip_graph_siglip(clip_ctx * ctx, const clip_image_f32 & img) : clip_graph(ctx, img) {}
|
|
ggml_cgraph * build() override;
|
|
};
|
|
|
|
struct clip_graph_gemma4v : clip_graph {
|
|
clip_graph_gemma4v(clip_ctx * ctx, const clip_image_f32 & img) : clip_graph(ctx, img) {}
|
|
ggml_cgraph * build() override;
|
|
ggml_tensor * build_mm(ggml_tensor * w, ggml_tensor * x) const override;
|
|
bool support_batch() const override { return true; }
|
|
};
|
|
|
|
struct clip_graph_gemma4uv : clip_graph {
|
|
clip_graph_gemma4uv(clip_ctx * ctx, const clip_image_f32 & img) : clip_graph(ctx, img) {}
|
|
ggml_cgraph * build() override;
|
|
};
|
|
|
|
struct clip_graph_pixtral : clip_graph {
|
|
clip_graph_pixtral(clip_ctx * ctx, const clip_image_f32 & img) : clip_graph(ctx, img) {}
|
|
ggml_cgraph * build() override;
|
|
};
|
|
|
|
struct clip_graph_qwen2vl : clip_graph {
|
|
clip_graph_qwen2vl(clip_ctx * ctx, const clip_image_f32 & img) : clip_graph(ctx, img) {}
|
|
ggml_cgraph * build() override;
|
|
ggml_tensor * build_inp_with_temporal_merge();
|
|
};
|
|
|
|
struct clip_graph_qwen3vl : clip_graph_qwen2vl {
|
|
clip_graph_qwen3vl(clip_ctx * ctx, const clip_image_f32 & img) : clip_graph_qwen2vl(ctx, img) {}
|
|
ggml_cgraph * build() override;
|
|
};
|
|
|
|
struct clip_graph_minimax_m3 : clip_graph {
|
|
clip_graph_minimax_m3(clip_ctx * ctx, const clip_image_f32 & img) : clip_graph(ctx, img) {}
|
|
ggml_cgraph * build() override;
|
|
ggml_tensor * apply_rope(ggml_tensor * x, ggml_tensor * pos_h, ggml_tensor * pos_w);
|
|
};
|
|
|
|
struct clip_graph_mimovl : clip_graph {
|
|
clip_graph_mimovl(clip_ctx * ctx, const clip_image_f32 & img) : clip_graph(ctx, img) {}
|
|
ggml_cgraph * build() override;
|
|
// Force F32 mat-mul accumulation to avoid F16 overflow in the FFN down-proj
|
|
// when the mmproj is stored in F16 (the source weights are BF16; downcasting
|
|
// to F16 reduces dynamic range below the SwiGLU output magnitude on the last few layers).
|
|
ggml_tensor * build_mm(ggml_tensor * w, ggml_tensor * x) const override;
|
|
};
|
|
|
|
struct clip_graph_step3vl : clip_graph {
|
|
clip_graph_step3vl(clip_ctx * ctx, const clip_image_f32 & img) : clip_graph(ctx, img) {}
|
|
ggml_cgraph * build() override;
|
|
};
|
|
|
|
struct clip_graph_youtuvl : clip_graph {
|
|
clip_graph_youtuvl(clip_ctx * ctx, const clip_image_f32 & img) : clip_graph(ctx, img) {}
|
|
ggml_cgraph * build() override;
|
|
};
|
|
|
|
struct clip_graph_yasa2 : clip_graph {
|
|
clip_graph_yasa2(clip_ctx * ctx, const clip_image_f32 & img) : clip_graph(ctx, img) {}
|
|
ggml_cgraph * build() override;
|
|
|
|
ggml_tensor * layer_norm_channels(ggml_tensor * inp, ggml_tensor * w, ggml_tensor * b, float eps = 1e-6f);
|
|
ggml_tensor * convnext_grn(ggml_tensor * inp, ggml_tensor * w, ggml_tensor * b);
|
|
};
|
|
|
|
struct clip_graph_minicpmv : clip_graph {
|
|
clip_graph_minicpmv(clip_ctx * ctx, const clip_image_f32 & img) : clip_graph(ctx, img) {}
|
|
ggml_cgraph * build() override;
|
|
};
|
|
|
|
struct clip_graph_minicpmv4_6 : clip_graph {
|
|
clip_graph_minicpmv4_6(clip_ctx * ctx, const clip_image_f32 & img) : clip_graph(ctx, img) {}
|
|
ggml_cgraph * build() override;
|
|
};
|
|
|
|
struct clip_graph_internvl : clip_graph {
|
|
clip_graph_internvl(clip_ctx * ctx, const clip_image_f32 & img) : clip_graph(ctx, img) {}
|
|
ggml_cgraph * build() override;
|
|
bool support_batch() const override { return true; }
|
|
};
|
|
|
|
struct clip_graph_nemotron_v2_vl : clip_graph {
|
|
clip_graph_nemotron_v2_vl(clip_ctx * ctx, const clip_image_f32 & img) : clip_graph(ctx, img) {}
|
|
ggml_cgraph * build() override;
|
|
};
|
|
|
|
struct clip_graph_llama4 : clip_graph {
|
|
clip_graph_llama4(clip_ctx * ctx, const clip_image_f32 & img) : clip_graph(ctx, img) {}
|
|
ggml_cgraph * build() override;
|
|
};
|
|
|
|
struct clip_graph_kimivl : clip_graph {
|
|
clip_graph_kimivl(clip_ctx * ctx, const clip_image_f32 & img) : clip_graph(ctx, img) {}
|
|
ggml_cgraph * build() override;
|
|
};
|
|
|
|
struct clip_graph_paddleocr : clip_graph {
|
|
clip_graph_paddleocr(clip_ctx * ctx, const clip_image_f32 & img) : clip_graph(ctx, img) {}
|
|
ggml_cgraph * build() override;
|
|
};
|
|
|
|
struct clip_graph_dotsocr : clip_graph {
|
|
clip_graph_dotsocr(clip_ctx * ctx, const clip_image_f32 & img) : clip_graph(ctx, img) {}
|
|
ggml_cgraph * build() override;
|
|
};
|
|
|
|
struct clip_graph_cogvlm : clip_graph {
|
|
clip_graph_cogvlm(clip_ctx * ctx, const clip_image_f32 & img) : clip_graph(ctx, img) {}
|
|
ggml_cgraph * build() override;
|
|
};
|
|
|
|
struct clip_graph_llava : clip_graph {
|
|
clip_graph_llava(clip_ctx * ctx, const clip_image_f32 & img) : clip_graph(ctx, img) {}
|
|
ggml_cgraph * build() override;
|
|
};
|
|
|
|
struct clip_graph_whisper_enc : clip_graph {
|
|
clip_graph_whisper_enc(clip_ctx * ctx, const clip_image_f32 & img) : clip_graph(ctx, img) {}
|
|
ggml_cgraph * build() override;
|
|
};
|
|
|
|
struct clip_graph_deepseekocr : clip_graph {
|
|
clip_graph_deepseekocr(clip_ctx * ctx, const clip_image_f32 & img) : clip_graph(ctx, img) {}
|
|
ggml_cgraph * build() override;
|
|
ggml_tensor * build_sam(ggml_tensor * inp); // build the SAM model
|
|
bool support_batch() const override { return true; }
|
|
};
|
|
|
|
struct clip_graph_deepseekocr2 : clip_graph_deepseekocr {
|
|
clip_graph_deepseekocr2(clip_ctx * ctx, const clip_image_f32 & img) : clip_graph_deepseekocr(ctx, img) {}
|
|
ggml_cgraph * build() override; // reuses build_sam() from base
|
|
bool support_batch() const override { return true; }
|
|
};
|
|
|
|
struct clip_graph_conformer : clip_graph {
|
|
clip_graph_conformer(clip_ctx * ctx, const clip_image_f32 & img) : clip_graph(ctx, img) {}
|
|
ggml_cgraph * build() override;
|
|
};
|
|
|
|
struct clip_graph_granite_speech : clip_graph {
|
|
clip_graph_granite_speech(clip_ctx * ctx, const clip_image_f32 & img) : clip_graph(ctx, img) {}
|
|
ggml_cgraph * build() override;
|
|
};
|
|
|
|
struct clip_graph_gemma4a : clip_graph {
|
|
clip_graph_gemma4a(clip_ctx * ctx, const clip_image_f32 & img) : clip_graph(ctx, img) {}
|
|
ggml_cgraph * build() override;
|
|
ggml_tensor * build_mm(ggml_tensor * w, ggml_tensor * x) const override;
|
|
};
|
|
|
|
struct clip_graph_gemma4ua : clip_graph {
|
|
clip_graph_gemma4ua(clip_ctx * ctx, const clip_image_f32 & img) : clip_graph(ctx, img) {}
|
|
ggml_cgraph * build() override;
|
|
};
|
|
|
|
struct clip_graph_glm4v : clip_graph {
|
|
clip_graph_glm4v(clip_ctx * ctx, const clip_image_f32 & img) : clip_graph(ctx, img) {}
|
|
ggml_cgraph * build() override;
|
|
};
|
|
|
|
struct clip_graph_hunyuanvl : clip_graph {
|
|
clip_graph_hunyuanvl(clip_ctx * ctx, const clip_image_f32 & img) : clip_graph(ctx, img) {}
|
|
ggml_cgraph * build() override;
|
|
};
|
|
|
|
struct clip_graph_mobilenetv5 : clip_graph {
|
|
clip_graph_mobilenetv5(clip_ctx * ctx, const clip_image_f32 & img) : clip_graph(ctx, img) {}
|
|
ggml_cgraph * build() override;
|
|
|
|
ggml_tensor * rms_norm_2d(
|
|
ggml_tensor * inp,
|
|
ggml_tensor * weight,
|
|
float eps = 1e-6f);
|
|
|
|
ggml_tensor* pad_same_2d(
|
|
ggml_tensor* inp,
|
|
int kernel_h,
|
|
int kernel_w,
|
|
int stride_h,
|
|
int stride_w,
|
|
int dilation_h = 1,
|
|
int dilation_w = 1);
|
|
|
|
ggml_tensor * build_edge_residual(
|
|
ggml_tensor * inp,
|
|
const mobilenetv5_block & block,
|
|
int stride);
|
|
|
|
ggml_tensor * build_inverted_residual(
|
|
ggml_tensor * inp,
|
|
const mobilenetv5_block & block,
|
|
int stride);
|
|
|
|
ggml_tensor * build_mobilenet_attn(
|
|
ggml_tensor * inp,
|
|
const mobilenetv5_block & block);
|
|
};
|
|
|
|
struct clip_graph_qwen3a : clip_graph {
|
|
clip_graph_qwen3a(clip_ctx * ctx, const clip_image_f32 & img) : clip_graph(ctx, img) {}
|
|
ggml_cgraph * build() override;
|
|
};
|
|
|
|
struct clip_graph_mimo_audio : clip_graph {
|
|
clip_graph_mimo_audio(clip_ctx * ctx, const clip_image_f32 & img) : clip_graph(ctx, img) {}
|
|
ggml_cgraph * build() override;
|
|
};
|
|
|
|
struct clip_graph_qwen3tts_spkenc : clip_graph {
|
|
clip_graph_qwen3tts_spkenc(clip_ctx * ctx, const clip_image_f32 & img) : clip_graph(ctx, img) {}
|
|
ggml_cgraph * build() override;
|
|
|
|
ggml_tensor * conv1d_same(ggml_tensor * x, ggml_tensor * w, ggml_tensor * b, int dilation) const;
|
|
ggml_tensor * res2net(ggml_tensor * x, const clip_layer & layer, int dilation, int scale) const;
|
|
ggml_tensor * se_block(ggml_tensor * x, const clip_layer & layer) const;
|
|
ggml_tensor * se_res2net_block(ggml_tensor * x, const clip_layer & layer, int dilation, int scale) const;
|
|
ggml_tensor * attentive_stats_pool(ggml_tensor * x) const;
|
|
};
|
|
|
|
struct clip_graph_qwen3tts_gen : clip_graph {
|
|
clip_graph_qwen3tts_gen(clip_ctx * ctx, const clip_image_f32 & img, clip_gen_process_type gen_process, int top_k, float top_p)
|
|
: clip_graph(ctx, img), gen_process(gen_process), top_k(top_k), top_p(top_p) {}
|
|
ggml_cgraph * build() override;
|
|
|
|
// which sub-graph build() constructs, fixed at graph-build time
|
|
clip_gen_process_type gen_process;
|
|
|
|
// sampling params, fixed at graph-build time (GEN_CODE only)
|
|
int top_k;
|
|
float top_p;
|
|
|
|
//
|
|
// code_gen: backbone hidden state + sampled code0 -> 16 RVQ codes
|
|
// MTP-style code predictor, one token per codebook
|
|
//
|
|
struct code_gen : clip_graph {
|
|
code_gen(const clip_graph & parent, int top_k, float top_p)
|
|
: clip_graph(parent), top_k(top_k), top_p(top_p) {}
|
|
ggml_cgraph * build() override { GGML_ABORT("call prefill()/step() instead"); }
|
|
|
|
int top_k;
|
|
float top_p;
|
|
|
|
ggml_tensor * cache_set(ggml_tensor * cache, int row_idx, ggml_tensor * value) const;
|
|
ggml_tensor * do_sampling(ggml_tensor * logits, ggml_tensor * inp_rand) const;
|
|
|
|
ggml_tensor * const_i32(ggml_tensor * anchor, float value) const;
|
|
ggml_tensor * causal_mask_row(int64_t n_kv_pad, int pos) const;
|
|
ggml_tensor * project_in(ggml_tensor * cur) const;
|
|
|
|
ggml_tensor * layer_forward(
|
|
ggml_tensor * cur,
|
|
const clip_layer & layer,
|
|
ggml_tensor * inp_pos,
|
|
ggml_tensor * kq_mask,
|
|
ggml_tensor *& k_cache_layer,
|
|
ggml_tensor *& v_cache_layer,
|
|
int64_t n_kv_pad,
|
|
int pos,
|
|
int il) const;
|
|
|
|
void prefill(
|
|
std::vector<ggml_tensor *> & k_cache,
|
|
std::vector<ggml_tensor *> & v_cache,
|
|
ggml_tensor *& out_code_cache,
|
|
ggml_tensor * h_state,
|
|
ggml_tensor * code0_embd,
|
|
ggml_tensor * inp_rand) const;
|
|
|
|
ggml_tensor * step(
|
|
std::vector<ggml_tensor *> & k_cache,
|
|
std::vector<ggml_tensor *> & v_cache,
|
|
ggml_tensor * out_code_cache,
|
|
ggml_tensor * inp_rand,
|
|
int step_idx) const;
|
|
};
|
|
|
|
//
|
|
// code2wav: RVQ codes -> raw PCM (quantizer + pre_conv + pre_transformer + upsample + DAC).
|
|
//
|
|
struct code2wav : clip_graph {
|
|
code2wav(const clip_graph & parent) : clip_graph(parent) {}
|
|
ggml_cgraph * build() override { GGML_ABORT("call decode() instead"); }
|
|
|
|
// state_in: previous call's persisted state, by slot name (see list_c2w_state_slots())
|
|
std::map<std::string, ggml_tensor *> state_in;
|
|
// state_out: this call's state to persist, added to the graph outputs by build()
|
|
mutable std::vector<std::pair<std::string, ggml_tensor *>> state_out;
|
|
|
|
// stateful conv ops: read/update their state via state_in/state_out[state_name]
|
|
ggml_tensor * causal_conv1d(ggml_tensor * x, ggml_tensor * w, ggml_tensor * b, int dilation, const std::string & state_name) const;
|
|
ggml_tensor * causal_conv1d_dw(ggml_tensor * x, ggml_tensor * w, ggml_tensor * b, const std::string & state_name) const;
|
|
ggml_tensor * causal_conv_transpose1d(ggml_tensor * x, ggml_tensor * w, ggml_tensor * b, int stride, const std::string & state_name) const;
|
|
ggml_tensor * snake(ggml_tensor * x, ggml_tensor * alpha, ggml_tensor * beta) const;
|
|
|
|
ggml_tensor * quant_decode(ggml_tensor * inp_codes) const;
|
|
ggml_tensor * tfm_layer_forward(ggml_tensor * cur, const clip_layer & layer, int il) const;
|
|
ggml_tensor * convnext_block(ggml_tensor * x, const clip_code2wav::upsample_block & blk, const std::string & state_prefix) const;
|
|
ggml_tensor * dac_res_unit(ggml_tensor * x, const clip_code2wav::dac_res & res, int dilation, const std::string & state_name) const;
|
|
|
|
// inp_codes [1, n_codes] I32 -> this frame's audio samples [n_samples] F32, clamped to [-1, 1]
|
|
ggml_tensor * decode(ggml_tensor * inp_codes) const;
|
|
};
|
|
};
|
|
|
|
// one persisted state buffer used by code2wav, see qwen3tts-gen.cpp
|
|
struct c2w_state_slot {
|
|
std::string name;
|
|
int64_t ne0;
|
|
int64_t ne1;
|
|
};
|
|
std::vector<c2w_state_slot> list_c2w_state_slots(const clip_hparams & hparams, const clip_model & model);
|
|
|
|
struct clip_graph_kimik25 : clip_graph {
|
|
clip_graph_kimik25(clip_ctx * ctx, const clip_image_f32 & img) : clip_graph(ctx, img) {}
|
|
ggml_cgraph * build() override;
|
|
|
|
ggml_tensor * resize_position_embeddings_3d(uint32_t interpolation_mode);
|
|
};
|
|
|
|
struct clip_graph_parakeet : clip_graph {
|
|
clip_graph_parakeet(clip_ctx * ctx, const clip_image_f32 & img) : clip_graph(ctx, img) {}
|
|
ggml_cgraph * build() override;
|
|
};
|
|
|
|
struct clip_graph_exaone4_5 : clip_graph {
|
|
clip_graph_exaone4_5(clip_ctx * ctx, const clip_image_f32 & img) : clip_graph(ctx, img) {}
|
|
ggml_cgraph * build() override;
|
|
};
|
|
|
|
struct clip_graph_granite4_vision : clip_graph {
|
|
clip_graph_granite4_vision(clip_ctx * ctx, const clip_image_f32 & img)
|
|
: clip_graph(ctx, img),
|
|
add_newline(img.add_newline) {}
|
|
|
|
ggml_cgraph * build() override;
|
|
|
|
private:
|
|
// The graph is per-tile since only batch-size 1 is supported in clip. As
|
|
// such, this value is set at construct time based on the tile that will be
|
|
// encoded, then used during build to determine how to handle newlines.
|
|
const bool add_newline;
|
|
|
|
ggml_tensor * gather(ggml_tensor * src, const std::string & name, int idx_len);
|
|
ggml_tensor * interp_down(ggml_tensor * src, int side, int new_side);
|
|
ggml_tensor * build_block(const qf_block & blk, ggml_tensor * h, int bid,
|
|
int spatial_offset, int image_side, int window_side,
|
|
int query_side, float qformer_eps);
|
|
|
|
ggml_tensor * build_newline_row(ggml_context * ctx0);
|
|
ggml_tensor * append_rowwise_newlines(ggml_context * ctx0, ggml_tensor * tile_output);
|
|
};
|
|
|
|
struct clip_graph_muse_glimmer : clip_graph {
|
|
clip_graph_muse_glimmer(clip_ctx * ctx, const clip_image_f32 & img) : clip_graph(ctx, img) {}
|
|
ggml_cgraph * build() override;
|
|
};
|