Not a member of Pastebin yet?
Sign Up,
it unlocks many cool features!
- RuntimeError: Expected to mark a variable ready only once. This error is caused by one of the following reasons: 1) Use of a module parameter outside the `forward` function. Please make sure model paramet
- ers are not shared across multiple concurrent forward-backward passes. or try to use _set_static_graph() as a workaround if this module graph does not change during training loop.2) Reused parameters in m
- ultiple reentrant backward passes. For example, if you use multiple `checkpoint` functions to wrap the same part of your model, it would result in the same set of parameters been used by different reentra
- nt backward passes multiple times, and hence marking a variable ready multiple times. DDP does not support such use cases in default. You can try to use _set_static_graph() as a workaround if your module
- graph does not change over iterations.
- Parameter at index 166 with name model.layers.27.mlp_norm.weight has been marked as ready twice. This means that multiple autograd engine hooks have fired for this particular parameter during this iterat
- ion.
- [rank0]: Traceback (most recent call last):
- [rank0]: File "/workspace/scoring_train_modernbert_custom.py", line 355, in <module>
- [rank0]: train_result = trainer.train()
- [rank0]: ^^^^^^^^^^^^^^^
- [rank0]: File "/workspace/pi_scorer_trainer.py", line 2507, in train
- [rank0]: return inner_training_loop(
- [rank0]: ^^^^^^^^^^^^^^^^^^^^
- [rank0]: File "/workspace/pi_scorer_trainer.py", line 2955, in _inner_training_loop
- [rank0]: tr_loss_step = self.training_step(
- [rank0]: ^^^^^^^^^^^^^^^^^^^
- [rank0]: File "/workspace/pi_scorer_trainer.py", line 4392, in training_step
- [rank0]: self.accelerator.backward(loss, **kwargs)
- [rank0]: File "/usr/local/lib/python3.11/dist-packages/accelerate/accelerator.py", line 2246, in backward
- [rank0]: loss.backward(**kwargs)
- [rank0]: File "/usr/local/lib/python3.11/dist-packages/torch/_tensor.py", line 521, in backward
- [rank0]: torch.autograd.backward(
- [rank0]: File "/usr/local/lib/python3.11/dist-packages/torch/autograd/__init__.py", line 289, in backward
- [rank0]: _engine_run_backward(
- [rank0]: File "/usr/local/lib/python3.11/dist-packages/torch/autograd/graph.py", line 769, in _engine_run_backward
- [rank0]: return Variable._execution_engine.run_backward( # Calls into the C++ engine to run the backward pass
- [rank0]: ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
- [rank0]: File "/usr/local/lib/python3.11/dist-packages/torch/autograd/function.py", line 306, in apply
- [rank0]: return user_fn(self, *args)
- [rank0]: ^^^^^^^^^^^^^^^^^^^^
- [rank0]: File "/usr/local/lib/python3.11/dist-packages/torch/utils/checkpoint.py", line 313, in backward
- [rank0]: torch.autograd.backward(outputs_with_grad, args_with_grad)
- [rank0]: File "/usr/local/lib/python3.11/dist-packages/torch/autograd/__init__.py", line 289, in backward
- [rank0]: _engine_run_backward(
- [rank0]: File "/usr/local/lib/python3.11/dist-packages/torch/autograd/graph.py", line 769, in _engine_run_backward
- [rank0]: return Variable._execution_engine.run_backward( # Calls into the C++ engine to run the backward pass
- [rank0]: ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
- [rank0]: RuntimeError: Expected to mark a variable ready only once. This error is caused by one of the following reasons: 1) Use of a module parameter outside the `forward` function. Please make sure mode
- l parameters are not shared across multiple concurrent forward-backward passes. or try to use _set_static_graph() as a workaround if this module graph does not change during training loop.2) Reused parame
- ters in multiple reentrant backward passes. For example, if you use multiple `checkpoint` functions to wrap the same part of your model, it would result in the same set of parameters been used by differen
- t reentrant backward passes multiple times, and hence marking a variable ready multiple times. DDP does not support such use cases in default. You can try to use _set_static_graph() as a workaround if you
- r module graph does not change over iterations.
Advertisement
Add Comment
Please, Sign In to add comment