Pytorch loss implementation
Implement loss functions in PyTorch with proper tensor operations.From its SKILL.md
npx -y skills add cxcscmu/SkillLearnBench --skill pytorch-loss-implementationAssembled from the repository path, not quoted from the project. Check it against their README if it does not work.
SKILL.md
2.2 KB, 572 tokens by cl100k_base, as published. Nobody here has run it
PyTorch Loss Implementation
Key Concepts
1. Tensor Operations
import torch
# Sigmoid function
sigmoid_output = torch.sigmoid(input_tensor)
# Log function
log_output = torch.log(input_tensor)
# Mean reduction
mean_loss = loss.mean()
# Sum reduction
sum_loss = loss.sum()
2. Batch Processing
# Batch dimension handling
batch_size = tensor.shape[0]
x = tensor[:batch_size//2] # First half
y = tensor[batch_size//2:] # Second half
# Ensure same device and dtype
tensor = tensor.to(device=model.device, dtype=torch.float32)
3. Numerical Stability
Log-Sigmoid Stability
# Avoid: log(sigmoid(x)) can cause numerical issues
# Instead use:
stable_loss = torch.nn.functional.logsigmoid(x)
# Or manually:
loss = -torch.log(torch.sigmoid(x) + 1e-10)
Handling Small Values
# Add epsilon to avoid log(0)
safe_log = torch.log(value + 1e-8)
# Clamp to valid range
clamped = torch.clamp(value, min=1e-10, max=1.0)
4. Device Handling
# Ensure all tensors on same device
device = torch.device('cuda' if torch.cuda.is_available() else 'cpu')
tensor = tensor.to(device)
# Or get from model
device = next(model.parameters()).device
tensor = tensor.to(device)
Loss Function Pattern
def compute_loss(logits, labels, temperature=1.0, margin=0.5):
# 1. Normalize/compute rewards
rewards = logits / (sequence_length + 1e-8)
# 2. Compute differences
diff = temperature * rewards[:n//2] - temperature * rewards[n//2:] - margin
# 3. Apply objective
loss_per_pair = -torch.log(torch.sigmoid(diff) + 1e-10)
# 4. Reduce
loss = loss_per_pair.mean()
return loss
Debugging Tips
- Check tensor shapes at each step
- Use .detach() for inspecting values without affecting gradients
- Verify numerical stability with small inputs
- Test gradients with
loss.backward() - Print intermediate values for debugging
Performance Tips
- Use in-place operations where safe:
tensor.log_() - Avoid unnecessary cloning/copying
- Batch operations are faster than loops
- Use PyTorch functions over custom loops
What ships with it
Read from the repository
Just SKILL.md. No reference files, no scripts.