Parameter-efficient fine-tuning methods adapt generic pretrained models to different downstream tasks by tuning a small set of the network parameters. Recent methods adopt a semi-supervised approach to allocate a fixed parametric budget based on a weight sensitivity metric. The structured component assigns low-rank adapters to layers with the most sensitive weights. The unstructured component allocates the remaining updatable weights individually. An approximate definition of sensitivity is used in practice because the exact form is impossible to compute. We dissect the approximation and show that it favors layers with low standard deviation and is less reliable when used on individual weights. We perform an adversarial attack to highlight the bias and introduce a more robust version of approximate sensitivity. We also propose using sensitivity aggregated at the layer level to adapt the size of low-rank adapters before training, eliminating the need for unstructured tuning. This structured approach respects the predefined parametric budget and optimizes the assignment of adapters. We evaluate the proposed approach using the 19 datasets of the VTAB-1k benchmark with four parametric budgets.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Parameter-Efficient Fine Tuning with Debiased Sensitivity and Constrained Low-Rank Adaptation

  • Tom Pégeot,
  • Inna Kucher,
  • Adrian Popescu,
  • Bertrand Delezoide

摘要

Parameter-efficient fine-tuning methods adapt generic pretrained models to different downstream tasks by tuning a small set of the network parameters. Recent methods adopt a semi-supervised approach to allocate a fixed parametric budget based on a weight sensitivity metric. The structured component assigns low-rank adapters to layers with the most sensitive weights. The unstructured component allocates the remaining updatable weights individually. An approximate definition of sensitivity is used in practice because the exact form is impossible to compute. We dissect the approximation and show that it favors layers with low standard deviation and is less reliable when used on individual weights. We perform an adversarial attack to highlight the bias and introduce a more robust version of approximate sensitivity. We also propose using sensitivity aggregated at the layer level to adapt the size of low-rank adapters before training, eliminating the need for unstructured tuning. This structured approach respects the predefined parametric budget and optimizes the assignment of adapters. We evaluate the proposed approach using the 19 datasets of the VTAB-1k benchmark with four parametric budgets.