agentsclimarketplace

Run2 scala text processing

Skill cxcscmu/SkillLearnBench/skills/b2-self-feedback-gemini-3-flash-preview/python-scala-translation/run2_scala-text-processing

Robust text tokenization with regex and position tracking in Scala.From its SKILL.md

Install
npx -y skills add cxcscmu/SkillLearnBench --skill run2_scala-text-processing

Assembled from the repository path, not quoted from the project. Check it against their README if it does not work.

SKILL.md

0.8 KB, 184 tokens by cl100k_base, as published. Nobody here has run it

Text Tokenization in Scala

Scala String methods like split and replaceAll are available.

Regex Splitting

val words = text.split("\\s+").filter(_.nonEmpty)

Position Tracking

Find positions using indexOf.

def tokenizeWithPos(text: String): List[(String, Int, Int)] = {
  val words = text.split("\\s+").filter(_.nonEmpty)
  var curr = 0
  words.map { w =>
    val start = text.indexOf(w, curr)
    val end = start + w.length
    curr = end
    (w, start, end)
  }.toList
}

Character Sets

Scala's Set[Char] and exists for char-level operations. Using dropWhile and reverse.dropWhile for trimming custom character sets.

What ships with it

Read from the repository

Just SKILL.md. No reference files, no scripts.

Keep looking

Skills are one crate of 325,949. Ordering is by how many stacks a row turns up in, so the top of any crate is what has actually been picked rather than what has the most stars.