fix(caption): safetensors-only downloads, model load fixes, UI default, prefill tests

- Add use_safetensors=True to all 16 model from_pretrained calls to
  avoid downloading redundant .bin files alongside safetensors
- Add device property to JoyTag VisionModel so move_model can relocate
  it to CUDA (fixes 'ViT object has no attribute device')
- Fix Pix2Struct dtype mismatch by casting float inputs to model dtype
  while preserving integer tensor types
- Patch AutoConfig.register with exist_ok=True during Ovis loading to
  handle duplicate aimv2 registration on model reload
- Detect Qwen VL fine-tune architecture from config model_type instead
  of repo name, fixing ToriiGate and similar third-party fine-tunes
- Change UI default task from Short Caption to Normal Caption, and
  preserve it on model switch instead of resetting to Use Prompt
- Add dual-prefill testing across 5 VQA test methods using a shared
  _check_prefill helper
- Fix pre-existing ruff W605 in strip_think_xml_tags docstring
This commit is contained in:
CalamitousFelicitousness
2026-01-29 04:21:01 +00:00
parent 57659ab642
commit e2cdbe47fa
7 changed files with 88 additions and 14 deletions
+1
View File
@@ -69,6 +69,7 @@ def load(repo: str = None):
llava_model = LlavaForConditionalGeneration.from_pretrained(
repo,
torch_dtype=devices.dtype,
use_safetensors=True,
cache_dir=shared.opts.hfcache_dir,
**quant_args,
)