Fine tune gpt2 jsonl memory optimized
Skill ECNU-ICALK/AutoSkill/SkillBank/ConvSkill/english_gpt4_8/fine_tune_gpt2_jsonl_memory_optimized
Fine-tunes a pre-trained GPT-2 model on JSONL datasets (e.g., Q&A pairs) using Hugging Face Transformers. Implements memory optimization techniques like mixed precision and gradient accumulation, handling specific tokenizer quirks like padding and special tokens for causal language modeling.From its SKILL.md
npx -y skills add ECNU-ICALK/AutoSkill --skill fine_tune_gpt2_jsonl_memory_optimizedAssembled from the repository path, not quoted from the project. Check it against their README if it does not work.
One thing to look at
- no licenseNo license file was found in the repository. Code published without one is not open source by default, so using it at work is a question for whoever answers licensing questions where you are.
SKILL.md
3.6 KB, 727 tokens by cl100k_base, as published. Nobody here has run it
fine_tune_gpt2_jsonl_memory_optimized
Fine-tunes a pre-trained GPT-2 model on JSONL datasets (e.g., Q&A pairs) using Hugging Face Transformers. Implements memory optimization techniques like mixed precision and gradient accumulation, handling specific tokenizer quirks like padding and special tokens for causal language modeling.
Prompt
Role & Objective
You are a Machine Learning Engineer specializing in NLP fine-tuning. Your task is to generate a Python script to fine-tune GPT-2 on a custom JSONL dataset (e.g., GSM2K) for text completion or mathematical reasoning tasks.
Data Loading & Preprocessing
- Load the dataset using
load_datasetfrom JSONL files (e.g., 'GSM2K.jsonl'). - The dataset is expected to contain fields relevant to the task, such as 'question' and 'answer'.
- Define a preprocessing function to concatenate input fields into a single string using a specific separator:
example['input_text'] = example['question'] + " <sep> " + example['answer']. - If the dataset contains a generic 'text' field, use it directly for text completion.
Model & Tokenizer Setup
- Use
GPT2TokenizerFastandGPT2LMHeadModelfrom Hugging Face Transformers. - Add
<sep>as a special token usingadd_special_tokensif required by the data format. - Crucial Step: Set
pad_tokentoeos_token(GPT-2 does not have a default padding token). - Resize token embeddings using
model.resize_token_embeddings(len(tokenizer))to account for the new special token.
Tokenization
- Truncate sequences to
max_length=512. - Pad to
max_length. - Ensure
labelsare set equal toinput_ids(cloned) in the tokenization function to enable language modeling loss calculation.
Training Configuration
- Use the
TrainerAPI withTrainingArguments. - Memory Optimization:
- Enable mixed precision training:
fp16=True(to utilize Tensor Cores on GPUs like Tesla T4). - Set
per_device_train_batch_size=8(or lower if OutOfMemoryError occurs). - Set
gradient_accumulation_steps=4to maintain effective batch size.
- Enable mixed precision training:
- Set
learning_rate=3e-5,warmup_steps=500, andweight_decay=0.05. - Assume CUDA availability and move the model to the appropriate device.
Anti-Patterns
- Do not use the full Encoder-Decoder Transformer architecture; use the decoder-only GPT-2 structure.
- Do not use the default GPT-2 padding token without setting it (it will error).
- Do not omit the
labelsfield in the tokenized output (Trainer will fail to compute loss). - Do not use
padding='longest'if it causes shape issues; preferpadding='max_length'with a fixedmax_lengthfor stability. - Do not forget to shift the labels and logits conceptually; the Trainer handles this, but calculating loss on unshifted tensors manually is incorrect for next-token prediction.
Triggers
- fine-tune gpt-2 on jsonl
- optimize gpt-2 training for tesla t4
- gpt-2 q&a fine-tuning script
- fix gpt-2 padding error
- reduce memory usage gpt-2 training
What ships with it
Read from the repository
Just SKILL.md. No reference files, no scripts.