You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
What gets checkpointed is activations, not gradients: the forward keeps some of them and the backward recomputes the others. "Gradient checkpointing" comes from Chen et al. 2016 and OpenAI's gradient-checkpointing package, and transformers kept it. Most of the stack has since settled on "activation checkpointing":
So a user moving between accelerate's FSDP plugin, DeepSpeed and transformers meets the same technique under two names.
The timing fits: gradient_checkpointing_enable just gained every_n_layers and offload, and selective checkpointing (saving the attention output) is next (huggingface/trl#7194). Better to settle the name before more options pile up under the old one.
Proposal
TrainingArguments.activation_checkpointing / activation_checkpointing_kwargs, model.activation_checkpointing_enable() / _disable(), same arguments as today.
The gradient_checkpointing* names stay for one major as aliases with a FutureWarning, then go.
Internal names (supports_gradient_checkpointing, GradientCheckpointingLayer) are rarely user-facing: open question whether they move in the same major.
PEFT, accelerate and TRL switch in the same release cycle.
What gets checkpointed is activations, not gradients: the forward keeps some of them and the backward recomputes the others. "Gradient checkpointing" comes from Chen et al. 2016 and OpenAI's
gradient-checkpointingpackage, and transformers kept it. Most of the stack has since settled on "activation checkpointing":torch.utils.checkpoint,apply_activation_checkpointing)[activation_checkpoint]activation_checkpointingconfig sectionfsdp_activation_checkpointingjax.checkpoint/jax.rematgradient_checkpointingSo a user moving between accelerate's FSDP plugin, DeepSpeed and transformers meets the same technique under two names.
The timing fits:
gradient_checkpointing_enablejust gainedevery_n_layersandoffload, and selective checkpointing (saving the attention output) is next (huggingface/trl#7194). Better to settle the name before more options pile up under the old one.Proposal
TrainingArguments.activation_checkpointing/activation_checkpointing_kwargs,model.activation_checkpointing_enable()/_disable(), same arguments as today.gradient_checkpointing*names stay for one major as aliases with aFutureWarning, then go.supports_gradient_checkpointing,GradientCheckpointingLayer) are rarely user-facing: open question whether they move in the same major.Important
This issue is not open to external contribs