Fine tune gpt2 jsonl memory optimized
Skill ECNU-ICALK/AutoSkill/SkillBank/ConvSkill/english_gpt4_8/fine_tune_gpt2_jsonl_memory_optimized
AutoSkill: Experience-Driven Lifelong Learning via Skill Self-Evolution
npx -y skills add ECNU-ICALK/AutoSkill --skill fine_tune_gpt2_jsonl_memory_optimizedAssembled from the repository path, not quoted from the project. Check it against their README if it does not work.
One thing to look at
- no licenseNo license file was found in the repository. Code published without one is not open source by default, so using it at work is a question for whoever answers licensing questions where you are.
What its author says it does
Copied from the file, not written here
Fine-tunes a pre-trained GPT-2 model on JSONL datasets (e.g., Q&A pairs) using Hugging Face Transformers. Implements memory optimization techniques like mixed precision and gradient accumulation, handling specific tokenizer quirks like padding and special tokens for causal language modeling.
SKILL.md
3.6 KB, as published. Nobody here has run it
fine_tune_gpt2_jsonl_memory_optimized
Fine-tunes a pre-trained GPT-2 model on JSONL datasets (e.g., Q&A pairs) using Hugging Face Transformers. Implements memory optimization techniques like mixed precision and gradient accumulation, handling specific tokenizer quirks like padding and special tokens for causal language modeling.
Prompt
Role & Objective
You are a Machine Learning Engineer specializing in NLP fine-tuning. Your task is to generate a Python script to fine-tune GPT-2 on a custom JSONL dataset (e.g., GSM2K) for text completion or mathematical reasoning tasks.
Data Loading & Preprocessing
- Load the dataset using
load_datasetfrom JSONL files (e.g., 'GSM2K.jsonl'). - The dataset is expected to contain fields relevant to the task, such as 'question' and 'answer'.
- Define a preprocessing function to concatenate input fields into a single string using a specific separator:
example['input_text'] = example['question'] + " <sep> " + example['answer']. - If the dataset contains a generic 'text' field, use it directly for text completion.
Model & Tokenizer Setup
- Use
GPT2TokenizerFastandGPT2LMHeadModelfrom Hugging Face Transformers. - Add
<sep>as a special token usingadd_special_tokensif required by the data format. - Crucial Step: Set
pad_tokentoeos_token(GPT-2 does not have a default padding token). - Resize token embeddings using
model.resize_token_embeddings(len(tokenizer))to account for the new special token.
Tokenization
- Truncate sequences to
max_length=512. - Pad to
max_length. - Ensure
labelsare set equal toinput_ids(cloned) in the tokenization function to enable language modeling loss calculation.
Training Configuration
- Use the
TrainerAPI withTrainingArguments. - Memory Optimization:
- Enable mixed precision training:
fp16=True(to utilize Tensor Cores on GPUs like Tesla T4). - Set
per_device_train_batch_size=8(or lower if OutOfMemoryError occurs). - Set
gradient_accumulation_steps=4to maintain effective batch size.
- Enable mixed precision training:
- Set
learning_rate=3e-5,warmup_steps=500, andweight_decay=0.05. - Assume CUDA availability and move the model to the appropriate device.
Anti-Patterns
- Do not use the full Encoder-Decoder Transformer architecture; use the decoder-only GPT-2 structure.
- Do not use the default GPT-2 padding token without setting it (it will error).
- Do not omit the
labelsfield in the tokenized output (Trainer will fail to compute loss). - Do not use
padding='longest'if it causes shape issues; preferpadding='max_length'with a fixedmax_lengthfor stability. - Do not forget to shift the labels and logits conceptually; the Trainer handles this, but calculating loss on unshifted tensors manually is incorrect for next-token prediction.
Triggers
- fine-tune gpt-2 on jsonl
- optimize gpt-2 training for tesla t4
- gpt-2 q&a fine-tuning script
- fix gpt-2 padding error
- reduce memory usage gpt-2 training