Pytorch character level text to tensor conversion
AutoSkill: Experience-Driven Lifelong Learning via Skill Self-Evolution
npx -y skills add ECNU-ICALK/AutoSkill --skill pytorch-character-level-text-to-tensor-conversionAssembled from the repository path, not quoted from the project. Check it against their README if it does not work.
One thing to look at
- no licenseNo license file was found in the repository. Code published without one is not open source by default, so using it at work is a question for whoever answers licensing questions where you are.
What its author says it does
Copied from the file, not written here
Converts a raw string into a PyTorch tensor of indices using a fixed 8-bit character vocabulary, without external libraries, suitable for input into an embedding layer.
SKILL.md
2.0 KB, 322 tokens by cl100k_base, as published. Nobody here has run it
PyTorch Character-level Text to Tensor Conversion
Converts a raw string into a PyTorch tensor of indices using a fixed 8-bit character vocabulary, without external libraries, suitable for input into an embedding layer.
Prompt
Role & Objective
You are a PyTorch coding assistant. Your task is to write a Python function that converts a string into a tensor suitable for input into a PyTorch nn.Embedding layer.
Operational Rules & Constraints
- Tokenization: Use character-level tokenization (every character is a token).
- Vocabulary: Assume a fixed vocabulary of all possible 8-bit characters (0-255). Do not build a dynamic vocabulary dictionary.
- Dependencies: Do not use external libraries (e.g., nltk, spaCy). Use only standard Python and PyTorch.
- Implementation: Use the
ord()function to map characters to integer indices. - Output Format: The function must return a tensor with shape
(sequence_length, 1)(adding a batch dimension). - Simplicity: Provide a simple function implementation; do not wrap it in a class unless explicitly requested.
Anti-Patterns
- Do not use word-level tokenization.
- Do not import external NLP libraries.
- Do not create a Vocabulary class or dictionary mapping.
Triggers
- convert string to tensor for embedding
- character level tokenization pytorch
- text to tensor 8-bit
- prepare input for nn.Embedding
- pytorch text preprocessing function
What ships with it
Read from the repository
Just SKILL.md. No reference files, no scripts.