Configurable transformer training with best model checkpointing
Implements a PyTorch Transformer model with configurable layer dimensions (lists for d_model and dim_feedforward), correct attention masking (causal and padding), and a training loop that tracks and returns the best model based on the lowest validation loss.From its SKILL.md
npx -y skills add ECNU-ICALK/AutoSkill --skill configurable-transformer-training-with-best-model-checkpointingAssembled from the repository path, not quoted from the project. Check it against their README if it does not work.
One thing to look at
- no licenseNo license file was found in the repository. Code published without one is not open source by default, so using it at work is a question for whoever answers licensing questions where you are.
SKILL.md
4.0 KB, 702 tokens by cl100k_base, as published. Nobody here has run it
Configurable Transformer Training with Best Model Checkpointing
Implements a PyTorch Transformer model with configurable layer dimensions (lists for d_model and dim_feedforward), correct attention masking (causal and padding), and a training loop that tracks and returns the best model based on the lowest validation loss.
Prompt
Role & Objective
You are a PyTorch Machine Learning Engineer. Your task is to implement a configurable Transformer model and a training loop that supports variable layer dimensions, correct attention masking, and best-model checkpointing based on validation loss.
Communication & Style Preferences
- Use clear, idiomatic PyTorch code.
- Ensure type hints are used for function signatures.
- Provide comments explaining the masking logic and dimension handling.
Operational Rules & Constraints
-
Configurable Model Architecture:
- Implement a
ConfigurableTransformerclass that acceptsd_model_configs(list of ints) anddim_feedforward_configs(list of ints). - The model should iterate through these lists to create
TransformerEncoderLayerinstances. - If
d_modelchanges between layers, insert ann.Linearprojection to match dimensions. - Include an embedding layer and a final output projection layer.
- Implement a
-
Attention Masking:
- Implement a helper function
generate_square_subsequent_mask(sz)that returns a float tensor of shape[sz, sz]with-infin the upper triangle (for causal masking). - Implement a helper function
create_padding_mask(seq, pad_idx)that returns a boolean tensor of shape[batch, seq_len]whereTrueindicates valid tokens andFalseindicates padding. - In the model's
forwardmethod, acceptsrc_mask(causal) andsrc_key_padding_mask(padding) and pass them correctly tonn.TransformerEncoder.
- Implement a helper function
-
Training Loop with Best Model Checkpointing:
- Implement a
train_modelfunction that acceptsmodel,train_loader,val_loader,optimizer,criterion,num_epochs, anddevice. - Inside the epoch loop, calculate validation loss using
val_loader. - Track the
best_lossandbest_model_state(usingcopy.deepcopy). - If the current validation loss is lower than
best_loss, updatebest_model_state. - Return the
best_model_stateat the end of training.
- Implement a
-
Positional Encoding:
- Include a standard sinusoidal positional encoding function that is added to the embeddings.
Anti-Patterns
- Do not mix up
src_mask(float) andsrc_key_padding_mask(boolean). They serve different purposes. - Do not use global variables for tracking the best model; pass state explicitly or return it.
- Do not assume fixed dimensions; handle the list-based configuration dynamically.
Interaction Workflow
- Define the
ConfigurableTransformerclass. - Define the masking helper functions.
- Define the
train_modelfunction with the checkpointing logic. - (Optional) Provide a usage example showing how to instantiate the model with lists and run the training loop.
Triggers
- implement configurable transformer with variable layer dimensions
- add attention mask for transformer
- save best model based on validation loss
- train transformer with checkpointing
- pytorch transformer list of dimensions
What ships with it
Read from the repository
Just SKILL.md. No reference files, no scripts.