System Info
transformers version: 5.18.0.dev0
- Platform: macOS-15.2-arm64-arm-64bit
- Python version: 3.12.2
- Huggingface_hub version: 1.33.0
- Safetensors version: 0.8.0
- Accelerate version: 1.15.0
- Accelerate config: not found
- DeepSpeed version: not installed
- PyTorch version (accelerator?): 2.14.0 (NA)
- Using distributed or parallel set-up in script?:
Who can help?
@Cyrilvallez
Information
Tasks
Reproduction
import torch
from transformers import DeepseekOcr2Config, DeepseekOcr2Model
def report(title, model):
vt = model.vision_tower
print(f"--- {title}")
print("language_model :", model.language_model.config._attn_implementation)
print("vision_tower :", vt.config._attn_implementation)
print("vision_tower.sam_encoder :", vt.sam_encoder.config._attn_implementation)
print("vision_tower.vision_encoder :", vt.vision_encoder.config._attn_implementation)
with torch.device("meta"):
config = DeepseekOcr2Config(
text_config={"num_hidden_layers": 1, "mlp_layer_types": ["dense"]},
attn_implementation={"vision_config": "eager"},
)
model = DeepseekOcr2Model(config)
report("1. DeepseekOcr2Config(attn_implementation={'vision_config': 'eager'})", model)
with torch.device("meta"):
config = DeepseekOcr2Config(
text_config={"num_hidden_layers": 1, "mlp_layer_types": ["dense"]}
)
model = DeepseekOcr2Model(config)
model.set_attn_implementation({"vision_config": "eager"})
report("2. model.set_attn_implementation({'vision_config': 'eager'})", model)
Expected behavior
Setting attn_implementation at config construction vs via set_attn_implementation have different behaviors when the model have level-2 (and more) nested configs
Let's take DeepseekOcr2Model for example:
DeepseekOcr2Config
├── text_config
└── vision_config
├── sam_config
└── encoder_config
➡️ Setting it at config level
config = DeepseekOcr2Config(attn_implementation={"vision_config": "eager"})
--- DeepseekOcr2Config(attn_implementation={'vision_config': 'eager'})
language_model : sdpa
vision_tower : eager
vision_tower.sam_encoder : eager
vision_tower.vision_encoder : eager
➡️ but setting it via set_attn_implementation
model.set_attn_implementation({"vision_config": "eager"})
--- model.set_attn_implementation({'vision_config': 'eager'})
language_model : sdpa
vision_tower : eager
vision_tower.sam_encoder : sdpa
vision_tower.vision_encoder : sdpa
System Info
transformersversion: 5.18.0.dev0Who can help?
@Cyrilvallez
Information
Tasks
examplesfolder (such as GLUE/SQuAD, ...)Reproduction
Expected behavior
Setting
attn_implementationat config construction vs viaset_attn_implementationhave different behaviors when the model have level-2 (and more) nested configsLet's take
DeepseekOcr2Modelfor example:➡️ Setting it at config level
➡️ but setting it via
set_attn_implementation