Dimension Error After Prompt-tuning the Gemma2 model
Hi, I think it has something to do with the caching of generated results. This problem only arises with Gemma models, from what I experienced.
I had a similar problem a month ago. I was able to fix it by deploying the fix by @BenjaminB
github.com/huggingface/peft
#### FIX Prompt learning issue with 4d attention mask
`main` ← `BenjaminBossan:fix-prompt-tuning-4d-attention-mask`
opened 11:30AM - 27 Mar 25 UTC
BenjaminBossan
+100 -12
Resolves #2452 Some causal language models in transformers have 4d attention …masks at the input preparation stage. So far, we have assumed 2d attention masks, which results in an error in that case. This PR fixes the situation. My first attempt was to transform the 2d prefix attention mask (from the virtual tokens) into a 4d attention mask before concatenating them. However, this was error prone and I was unsure if my approach would generalize to other model architectures than the one tested (gemma), as it involved using private transformers methods (`model._prepare_4d_causal_attention_mask_with_cache_position`). The simpler approach was thus to just create a 2d attention mask and let the model handle it. The test suite has been extended to include a tiny gemma model. To prevent the test suite from ballooning, I removed another model. Specifically, this was GPT neox, which from HF download stats seems to be one of the least popular architectures from our test suite. I also extended the default parameters in `constants.py` for the different PEFT methods to support gemma. Unfortunately, some tests are failing with gemma. When they were unrelated to changes in this PR, I chose to just skip those tests, as I consider them out of scope for this PR.
It seems that it will be fixed in release 0.15.3. From now on, you can install the package from the main branch.
python -m pip install -U git+https://github.com/huggingface/peft.git@main