Distributed training
Skill thada2402/AutoResearchClaw/researchclaw/skills/builtin/tooling/distributed-training
Generate research papers autonomously by chatting with OpenClaw, using Python 3.11+, with a self-evolving framework and extensive test coverage.
npx -y skills add thada2402/AutoResearchClaw --skill distributed-trainingAssembled from the repository path, not quoted from the project. Check it against their README if it does not work.
One thing to look at
- 1 stars1 stars. Stars are a popularity signal and not a quality one, but at this level it is likely that nobody has read this closely except its author, and you would be relying on your own review.
What its author says it does
Copied from the file, not written here
Multi-GPU and distributed training patterns with PyTorch DDP. Use when scaling training across GPUs.
SKILL.md
0.8 KB, as published. Nobody here has run it
Distributed Training Best Practice
- Use DistributedDataParallel (DDP) over DataParallel for multi-GPU
- Initialize process group: dist.init_process_group(backend='nccl')
- Use DistributedSampler for data sharding
- Synchronize batch norm: nn.SyncBatchNorm.convert_sync_batchnorm()
- Only save checkpoint on rank 0
- Scale learning rate linearly with world size
- Use gradient accumulation for effectively larger batch sizes