Run2 scala text processing
Robust text tokenization with regex and position tracking in Scala.From its SKILL.md
npx -y skills add cxcscmu/SkillLearnBench --skill run2_scala-text-processingAssembled from the repository path, not quoted from the project. Check it against their README if it does not work.
SKILL.md
0.8 KB, 184 tokens by cl100k_base, as published. Nobody here has run it
Text Tokenization in Scala
Scala String methods like split and replaceAll are available.
Regex Splitting
val words = text.split("\\s+").filter(_.nonEmpty)
Position Tracking
Find positions using indexOf.
def tokenizeWithPos(text: String): List[(String, Int, Int)] = {
val words = text.split("\\s+").filter(_.nonEmpty)
var curr = 0
words.map { w =>
val start = text.indexOf(w, curr)
val end = start + w.length
curr = end
(w, start, end)
}.toList
}
Character Sets
Scala's Set[Char] and exists for char-level operations.
Using dropWhile and reverse.dropWhile for trimming custom character sets.
What ships with it
Read from the repository
Just SKILL.md. No reference files, no scripts.