Run3 read python tokenizer source
Use this skill to understand the exact Python source code in /root/Tokenizer.py before translating to Scala. This covers every class, method, regex pattern, enum value, default value, and metadata key that must be faithfully reproduced.From its SKILL.md
npx -y skills add cxcscmu/SkillLearnBench --skill run3_read-python-tokenizer-sourceAssembled from the repository path, not quoted from the project. Check it against their README if it does not work.
SKILL.md
1.6 KB, 310 tokens by cl100k_base, as published. Nobody here has run it
Reading and Understanding /root/Tokenizer.py
Step-by-step workflow
-
Read the file completely first:
cat /root/Tokenizer.py -
Key elements to extract and note:
TokenTypeenum: all member names and their string valuesTokendataclass/namedtuple: all fields, types, defaultsBaseTokenizerabstract class: interface methodsStringTokenizer: regex patterns, tokenize logic, what metadata it producesNumericTokenizer: regex patterns (integers, floats, scientific notation), metadata keysTemporalTokenizer: date/time regex patterns, metadataUniversalTokenizer: how it combines other tokenizers, fallback logicWhitespaceTokenizer: simple split logicTokenizerBuilder: builder pattern methods, what it configures- Free functions:
tokenize,tokenizeBatch,toToken,withMetadata— their exact signatures, parameters, defaults, return types
-
Pay special attention to:
- Every regex pattern string (copy exactly)
- Default parameter values
- Error handling (what happens with empty strings, None, invalid input)
- How metadata dictionaries are constructed
- How tokenizers are chained/composed in UniversalTokenizer
- Return types (lists, Optional values, etc.)
-
Document the exact API surface before writing any Scala code.
What ships with it
Read from the repository
Just SKILL.md. No reference files, no scripts.