Skip to content

[Fix] Make vision encoder DropPath initialization safe on meta devices - #1279

Open
cyyecao-lappland wants to merge 1 commit into
OpenGVLab:mainfrom
cyyecao-lappland:fix/issue-1254-meta-init
Open

cyyecao-lappland wants to merge 1 commit into
OpenGVLab:mainfrom
cyyecao-lappland:fix/issue-1254-meta-init

Conversation

@cyyecao-lappland

Copy link
Copy Markdown

Summary

Related to #1254.

The DropPath schedule inherits the default device from torch.linspace. Under a meta-device initialization context, extracting its Python scalar probabilities with .item() raises RuntimeError: Tensor.item() cannot be called on meta tensors.

  • Explicitly construct the schedule on CPU in both internvl_chat and internvl_chat_gpt_oss.
  • Preserve the existing schedule values and keep model parameters on the requested device.
  • Add small-config regression tests for both real vision implementations, without downloading pretrained weights. The tests import the vision modules under temporary package namespaces to avoid unrelated chat-backend initialization; model and dependency implementations are not mocked.

Thanks to @Liu-524 for identifying the explicit-CPU workaround in this comment.

Validation

Tested on Linux with Python 3.12.8, PyTorch 2.10.0+cu128, Transformers 5.1.0, timm 1.0.30, Accelerate 1.15.0, pytest 9.1.1, and an NVIDIA RTX 4090.

CUDA_VISIBLE_DEVICES=0 python -m pytest tests/test_intern_vit_initialization.py -q

The same regression suite produced:

  • Before the fix: 12 failed, 8 passed, with all failures caused by scalar extraction from meta tensors.
  • After the fix: 20 passed (CUDA cases ran, rather than being skipped).

Coverage includes CPU/meta parameter placement, one/four layers, zero/nonzero DropPath rates, exact schedule preservation, and small-config state-dict checkpoint round trips from meta initialization to CPU/CUDA with matching finite forward outputs on the same device.

git diff --check passes. Repository pre-commit checks on the changed files pass or have no applicable files, except flake8: both existing model files report W604 for docstring backticks at line 191. Running the same checks on the unmodified baseline reproduces those errors.

Scope and known limitations

This is a focused initialization fix, not a claim of full Transformers 5.x compatibility.

A separate tiny-model save_pretrained() / from_pretrained() probe moves past the original .item() failure after this patch, but still fails with AttributeError: 'InternVisionModel' object has no attribute 'all_tied_weights_keys' under Transformers 5.1.0. That loading issue and the separate generate/GenerationMixin reports are outside this patch.

No full pretrained InternVL checkpoint or end-to-end image/chat inference was tested. The passing checkpoint tests use PyTorch state-dict loading (assign=True), not Hugging Face from_pretrained().

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant