nam-2pir

Untitled

Jan 31st, 2025
100
0
Never
Not a member of Pastebin yet? Sign Up, it unlocks many cool features!
text 6.30 KB | None | 0 0
  1. RuntimeError: Expected to mark a variable ready only once. This error is caused by one of the following reasons: 1) Use of a module parameter outside the `forward` function. Please make sure model paramet
  2. ers are not shared across multiple concurrent forward-backward passes. or try to use _set_static_graph() as a workaround if this module graph does not change during training loop.2) Reused parameters in m
  3. ultiple reentrant backward passes. For example, if you use multiple `checkpoint` functions to wrap the same part of your model, it would result in the same set of parameters been used by different reentra
  4. nt backward passes multiple times, and hence marking a variable ready multiple times. DDP does not support such use cases in default. You can try to use _set_static_graph() as a workaround if your module
  5. graph does not change over iterations.
  6. Parameter at index 166 with name model.layers.27.mlp_norm.weight has been marked as ready twice. This means that multiple autograd engine hooks have fired for this particular parameter during this iterat
  7. ion.
  8. [rank0]: Traceback (most recent call last):
  9. [rank0]: File "/workspace/scoring_train_modernbert_custom.py", line 355, in <module>
  10. [rank0]: train_result = trainer.train()
  11. [rank0]: ^^^^^^^^^^^^^^^
  12. [rank0]: File "/workspace/pi_scorer_trainer.py", line 2507, in train
  13. [rank0]: return inner_training_loop(
  14. [rank0]: ^^^^^^^^^^^^^^^^^^^^
  15. [rank0]: File "/workspace/pi_scorer_trainer.py", line 2955, in _inner_training_loop
  16. [rank0]: tr_loss_step = self.training_step(
  17. [rank0]: ^^^^^^^^^^^^^^^^^^^
  18. [rank0]: File "/workspace/pi_scorer_trainer.py", line 4392, in training_step
  19. [rank0]: self.accelerator.backward(loss, **kwargs)
  20. [rank0]: File "/usr/local/lib/python3.11/dist-packages/accelerate/accelerator.py", line 2246, in backward
  21. [rank0]: loss.backward(**kwargs)
  22. [rank0]: File "/usr/local/lib/python3.11/dist-packages/torch/_tensor.py", line 521, in backward
  23. [rank0]: torch.autograd.backward(
  24. [rank0]: File "/usr/local/lib/python3.11/dist-packages/torch/autograd/__init__.py", line 289, in backward
  25. [rank0]: _engine_run_backward(
  26. [rank0]: File "/usr/local/lib/python3.11/dist-packages/torch/autograd/graph.py", line 769, in _engine_run_backward
  27. [rank0]: return Variable._execution_engine.run_backward( # Calls into the C++ engine to run the backward pass
  28. [rank0]: ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
  29. [rank0]: File "/usr/local/lib/python3.11/dist-packages/torch/autograd/function.py", line 306, in apply
  30. [rank0]: return user_fn(self, *args)
  31. [rank0]: ^^^^^^^^^^^^^^^^^^^^
  32. [rank0]: File "/usr/local/lib/python3.11/dist-packages/torch/utils/checkpoint.py", line 313, in backward
  33. [rank0]: torch.autograd.backward(outputs_with_grad, args_with_grad)
  34. [rank0]: File "/usr/local/lib/python3.11/dist-packages/torch/autograd/__init__.py", line 289, in backward
  35. [rank0]: _engine_run_backward(
  36. [rank0]: File "/usr/local/lib/python3.11/dist-packages/torch/autograd/graph.py", line 769, in _engine_run_backward
  37. [rank0]: return Variable._execution_engine.run_backward( # Calls into the C++ engine to run the backward pass
  38. [rank0]: ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
  39. [rank0]: RuntimeError: Expected to mark a variable ready only once. This error is caused by one of the following reasons: 1) Use of a module parameter outside the `forward` function. Please make sure mode
  40. l parameters are not shared across multiple concurrent forward-backward passes. or try to use _set_static_graph() as a workaround if this module graph does not change during training loop.2) Reused parame
  41. ters in multiple reentrant backward passes. For example, if you use multiple `checkpoint` functions to wrap the same part of your model, it would result in the same set of parameters been used by differen
  42. t reentrant backward passes multiple times, and hence marking a variable ready multiple times. DDP does not support such use cases in default. You can try to use _set_static_graph() as a workaround if you
  43. r module graph does not change over iterations.
Advertisement
Add Comment
Please, Sign In to add comment