ON PROGRAM RESTORATION FROM CHECKPOINTS SET

Authors

  • A. Y. Polyakov Author

Abstract

In paper two approaches to distributed programs restore problem from checkpoints set are described. Computation node wide algorithm of parent-child relationships and group/session assignement recreation at restore time is proposed. Also coordinated algorithm for process set restoration from several nodes/terminals is designed. Described algorightms are implemented in checkpointing package called extitDMTCP ( extitDistributed MultiThreaded CheckPointing).

Author Biography

  • A. Y. Polyakov
    Computer Systems Laboratory

Issue

Section

Mathematical Modelling