agentsclimarketplace

Write sequences

Skill FridrichMethod/awesome-skills/skills/write-sequences

Write biological sequences to files (FASTA, FASTQ, GenBank, EMBL) using Biopython Bio.SeqIO. Use when saving sequences, creating new sequence files, or outputting modified records.From its SKILL.md

Install
npx -y skills add FridrichMethod/awesome-skills --skill write-sequences

Assembled from the repository path, not quoted from the project. Check it against their README if it does not work.

2 things to look at

  • no licenseNo license file was found in the repository. Code published without one is not open source by default, so using it at work is a question for whoever answers licensing questions where you are.
  • 13 stars13 stars. Stars are a popularity signal and not a quality one, but at this level it is likely that nobody has read this closely except its author, and you would be relying on your own review.

SKILL.md

8.3 KB, ~2.1k tokens by cl100k_base, as published. Nobody here has run it

Version Compatibility

Reference examples tested with: BioPython 1.83+

Before using code patterns, verify installed versions match. If versions differ:

  • Python: pip show <package> then help(module.function) to check signatures

If code throws ImportError, AttributeError, or TypeError, introspect the installed package and adapt the example to match the actual API rather than retrying.

Write Sequences

"Write sequences to a file" -> Serialize SeqRecord objects into a formatted sequence file.

  • Python: SeqIO.write() (BioPython)
  • R: writeXStringSet() (Biostrings)

Governing Principle

The write is only as complete as the SeqRecord. Each format reads specific record fields and silently ignores the rest, so what survives a write is decided by which fields are populated before the call, not by the format string. FASTA serializes only id/description+seq; FASTQ additionally requires letter_annotations['phred_quality']; GenBank/EMBL additionally require annotations['molecule_type']. Populate the fields a format needs, or the write either drops data quietly (FASTA) or raises (FASTQ/GenBank).

Required Import

from Bio import SeqIO
from Bio.Seq import Seq
from Bio.SeqRecord import SeqRecord

Core Functions

SeqIO.write() - Write Records to File

SeqIO.write(records, 'output.fasta', 'fasta')
  • records - Single SeqRecord, list, or iterator of SeqRecords
  • handle - Filename (string) or open file handle
  • format - Lowercase output format string
  • Returns the number of records written (integer)

record.format() - Get Formatted String

formatted = record.format('fasta')

The FASTA Header Trap (id vs description)

FASTA output is built from record.description, NOT record.id. The writer compares the first whitespace token of description to id: if they match it writes description as-is; otherwise it prepends id + a space. When a record is parsed from FASTA, description already leads with the id token, so a round trip is faithful. But when id and description are set independently, a stale id-like token inside the description gets duplicated, and older BioPython releases dropped the id entirely instead of prepending it.

record.idrecord.descriptionHeader written
seq1seq1 kinase domain>seq1 kinase domain (clean: description leads with id)
seq1kinase domain>seq1 kinase domain (id auto-prepended)
seq1`` (empty)>seq1 (id used as fallback)
seq1gene7 kinase domain>seq1 gene7 kinase domain (stale id duplicated)

To control the header exactly and stay robust across versions, make description begin with id + a space: description=f'{rec_id} kinase domain'. The FASTA writer wraps the sequence at 60 characters per line by default.

Format Field Requirements

FormatStringRecord fields readHard requirement
FASTA'fasta'id/description, seqnone (header trap above)
FASTQ'fastq'seq, letter_annotationsphred quality scores
GenBank'genbank' / 'gb'seq, annotations, featuresmolecule_type
EMBL'embl'seq, annotations, featuresmolecule_type
Tab'tab'id, seqnone

Creating SeqRecord Objects

Goal: Construct in-memory records that carry the fields the target format requires.

Approach: Build a SeqRecord from a Seq plus id; add letter_annotations['phred_quality'] for FASTQ and annotations['molecule_type'] for GenBank/EMBL.

"Create a sequence record from scratch" -> Wrap a Seq in a SeqRecord with metadata.

  • Python: SeqRecord(Seq(...), id=...) (BioPython)
record = SeqRecord(Seq('ATGCGATCGATCG'), id='seq1', description='seq1 example sequence')

Code Patterns

Write Single or Multiple Records

records = [SeqRecord(Seq('ATGC'), id='seq1'), SeqRecord(Seq('GCTA'), id='seq2')]
count = SeqIO.write(records, 'output.fasta', 'fasta')

Write to a File Handle (and Append)

with open('output.fasta', 'w') as handle:
    SeqIO.write(records, handle, 'fasta')

with open('output.fasta', 'a') as handle:
    SeqIO.write(new_records, handle, 'fasta')

Write Modified Records via Generator

Goal: Transform sequences in memory and write the modified versions to a new file.

Approach: Parse input, map a transform over a generator, write the generator. Streaming avoids loading every record into RAM.

"Modify sequences and save" -> Parse records, transform each, write with SeqIO.write().

def uppercase_record(rec):
    return SeqRecord(rec.seq.upper(), id=rec.id, description=rec.description)

records = SeqIO.parse('input.fasta', 'fasta')
modified = (uppercase_record(rec) for rec in records)
SeqIO.write(modified, 'output.fasta', 'fasta')

Write FASTQ with Quality Scores

FASTQ requires letter_annotations['phred_quality'] as a list of ints. letter_annotations is length-locked to len(seq): assigning a list whose length differs from the sequence raises. Set the sequence first, then the quality list of matching length.

record = SeqRecord(Seq('ATGCGATCG'), id='read1')
record.letter_annotations['phred_quality'] = [40] * len(record.seq)
SeqIO.write(record, 'output.fastq', 'fastq')

Quality Encoding on Write (Phred vs Solexa)

When both phred_quality and solexa_quality keys are present, the writer uses Phred. Writing 'fastq-solexa' from a Phred-only record forces an on-the-fly lossy conversion (the scales diverge in the low-quality region) and emits a BiopythonWarning once any score reaches the high end (max quality >= ~62). For modern data, write plain 'fastq' (Sanger/Phred+33); only use 'fastq-solexa'/'fastq-illumina' when a tool explicitly demands that legacy encoding.

Write GenBank Format

GenBank and EMBL writing requires annotations['molecule_type'] (the alphabet that once carried this was removed in BioPython 1.78). Missing it raises on write.

record = SeqRecord(Seq('ATGCGATCGATCG'), id='SEQ001', name='example')
record.annotations['molecule_type'] = 'DNA'
record.annotations['topology'] = 'linear'
record.annotations['organism'] = 'Example organism'
SeqIO.write(record, 'output.gb', 'genbank')

Common Errors

SymptomCauseFix
Header has a duplicated or mangled idFASTA builds the header from description; it does not lead with id + spaceSet description=f'{rec.id} ...' or leave description empty to fall back to id
ValueError: No suitable quality scores found in letter_annotations of SeqRecord (id=...) on FASTQ writeRecord has no letter_annotations['phred_quality']Assign record.letter_annotations['phred_quality'] = [q]*len(seq)
TypeError: Any per-letter annotation should be a Python sequence ... of the same lengthQuality list length != len(seq) (annotations are length-locked)Set seq first, then a quality list of matching length
ValueError: missing molecule_type ... on GenBank/EMBL writeNo annotations['molecule_type'] since the 1.78 alphabet removalAdd record.annotations['molecule_type'] = 'DNA' (or 'RNA'/'protein')
BiopythonWarning: Data loss - max Solexa quality ...Writing 'fastq-solexa' from a high Phred-only record forces lossy conversionWrite plain 'fastq' unless a tool requires the Solexa encoding
TypeError passing a raw str/Seq to writeSeqIO.write expects SeqRecord(s)Wrap the sequence in a SeqRecord first
ValueError: Sequences must all be the same lengthPHYLIP/alignment format with unequal lengthsAlign, pad, or trim to equal length first

Related Skills

  • read-sequences - Read sequences before modifying and writing
  • format-conversion - Direct format conversion without intermediate processing
  • filter-sequences - Filter sequences before writing a subset
  • fastq-quality - Phred/Solexa encodings and quality-score handling
  • sequence-manipulation/seq-objects - Create SeqRecord objects to write
  • alignment-files/sam-bam-basics - For SAM/BAM output, use samtools/pysam

What ships with it: 2 files

4.8 KB alongside SKILL.md, 1 of them executable

examples/

Keep looking

Skills are one crate of 326,790. Ordering is by how many stacks a row turns up in, so the top of any crate is what has actually been picked rather than what has the most stars.